> Markdown version of [/videos/100496-kv-cache-is-not-about-speed-it-s-about-surviving-inference-costs](https://www.wearedevelopers.com/videos/100496-kv-cache-is-not-about-speed-it-s-about-surviving-inference-costs). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # KV Cache Is Not About Speed: It's About Surviving Inference Costs KV caching isn't about saving milliseconds on token generation. It's a survival strategy to prevent GPU memory exhaustion during the prefill phase and drastically slash LLM inference costs. - **Speakers:** [David vonThenen](https://www.wearedevelopers.com/@david-vonthenen) - **Event:** World Congress 2026 North America - **Published:** September 25, 2026 - **Duration:** 25:52 - **URL:** https://www.wearedevelopers.com/videos/100496-kv-cache-is-not-about-speed-it-s-about-surviving-inference-costs ## Access Playback and chapters for this video are available with a Free account. ## Related Moments - [Mitigating latency and memory walls with KV cache](https://www.wearedevelopers.com/videos/100488-point-ask-answer-building-vision-into-ai-on-constrained-hardware) (from "Point. Ask. Answer. Building Vision into AI on Constrained Hardware") - [Optimizing AI infrastructure costs by maximizing token caching](https://www.wearedevelopers.com/videos/100024-what-500-production-environments-taught-us-about-shipping-ai-agents) (from "What 500+ Production Environments Taught Us About Shipping AI Agents") - [Managing memory overhead by optimizing the KV cache](https://www.wearedevelopers.com/videos/100472-the-new-bottleneck-in-software-development) (from "The New Bottleneck in Software Development") - [Addressing caching pitfalls with specialized tools and managed services](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) (from "Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)") - [Optimizing unit economics through routing, caching, and compression](https://www.wearedevelopers.com/videos/100596-the-unit-economics-of-ai) (from "The Unit Economics of AI") - [Navigating the layers of the language model inference stack](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) (from "Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated") ## Related Articles - [Introducing Redis Agent Memory Server](https://www.wearedevelopers.com/magazine/699-introducing-redis-agent-memory-server) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [AI Eats the Verifiable First](https://www.wearedevelopers.com/magazine/765-ai-eats-the-verifiable-first) ## Related Jobs - [Principal Software Engineer, AI Inference Runtime](https://www.wearedevelopers.com/jobs/ext/2854958-principal-software-engineer-ai-inference-runtime) at **ARM** - [LLM Training Engineer](https://www.wearedevelopers.com/jobs/48420-llm-training-engineer) at **Sciforium** - [Staff Software Engineer, AI Inference Runtime](https://www.wearedevelopers.com/jobs/ext/3474466-staff-software-engineer-ai-inference-runtime) at **ARM** - [GPU Cluster Engineer, Systems & Platform](https://www.wearedevelopers.com/jobs/48411-gpu-cluster-engineer-systems-platform) at **Sciforium** - [Lead Software Engineer, Model Serving Platform](https://www.wearedevelopers.com/jobs/48413-lead-software-engineer-model-serving-platform) at **Sciforium** - [Distributed Training and Inference Engineer](https://www.wearedevelopers.com/jobs/48399-distributed-training-and-inference-engineer) at **Sciforium**