World Congress 2026 North America

KV Cache Is Not About Speed: It's About Surviving Inference Costs

September 23–25, 2026

World Congress 2026 North America

September 23–25, 2026 · San José, CA

Attend in person

Get tickets

Watch remotely

Watch live with Pro

Pro

Can’t make it to San José? Watch this session live with Pro. You also get:

  • All full videos, bookmarks, and playlists
  • World Congress livestreams
See pricing

What this session covers

Most teams think KV cache is about making inference faster. That’s true, but it’s not the point. KV cache is really about controlling memory, reducing recomputation, and keeping costs from spiraling as usage grows. As models get deployed at scale, the real bottleneck is no longer raw compute. It’s memory, bandwidth, and power. This session takes a step back to explain what KV cache does at a system level, why it matters for real workloads, and how approaches like vLLM, LMCache, and SGLang change how we think about scaling inference.

We’ll also connect this to a problem many teams are already seeing: confidently incorrect answers in agent systems. When cache behavior, context reuse, and routing aren’t designed well, systems don’t just get slower or more expensive. They get inconsistent. And that shows up as wrong answers with high confidence. This session will walk through these trade-offs using live demos, showing how different KV cache strategies impact cost, latency, and output quality in real time.

Related talks at this congress

Open session

World Congress 2026 North America

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization

Legare Kerrison, Cedric Clyburn

Legare Kerrison
Cedric Clyburn
Open session

World Congress 2026 North America

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

Context Engineering Kung Fu

Carl Lapierre

Tech Lead and AI Engineer at Osedea

Carl Lapierre
Open session

World Congress 2026 North America

Understanding LLM Architectures: Inside the Design of Modern Models

Jofia Jose Prakash

Enterprise AI Architect at American Chemical Society

Jofia Jose Prakash
All sessions at this congress