World Congress 2026 North America • Sep 25, 2026 • Session details

KV Cache Is Not About Speed: It's About Surviving Inference Costs

David vonThenen

KV caching isn't about saving milliseconds on token generation. It's a survival strategy to prevent GPU memory exhaustion during the prefill phase and drastically slash LLM inference costs.

KV Cache Is Not About Speed: It's About Surviving Inference Costs thumbnail

Checking access…

Playback and chapters load privately for Free videos.

Matching moments

3:31 min

Mitigating latency and memory walls with KV cache

Kavya Sri Chennoju Kavya Sri Chennoju · World Congress 2026 North America

4:26 min

Optimizing AI infrastructure costs by maximizing token caching

Liran Hason Liran Hason · World Congress 2026 Europe

2:34 min

Managing memory overhead by optimizing the KV cache

Aditya Jayaprakash Aditya Jayaprakash · World Congress 2026 North America

3:04 min

Addressing caching pitfalls with specialized tools and managed services

3:38 min

Optimizing unit economics through routing, caching, and compression

Ed Huang Ed Huang +4 · World Congress 2026 North America

2:00 min

Navigating the layers of the language model inference stack

Christin Pohl Christin Pohl · World Congress 2026 Europe