World Congress 2026 North America

KV Cache Is Not About Speed: It's About Surviving Inference Costs

September 24, 2026 15:30 – 16:00 · 30 min Stage 2

World Congress 2026 North America

September 23–25, 2026 · San José, CA

What this session covers

Most teams think KV cache is about making inference faster. That’s true, but it’s not the point. KV cache is really about controlling memory, reducing recomputation, and keeping costs from spiraling as usage grows. As models get deployed at scale, the real bottleneck is no longer raw compute. It’s memory, bandwidth, and power. This session takes a step back to explain what KV cache does at a system level, why it matters for real workloads, and how approaches like vLLM, LMCache, and SGLang change how we think about scaling inference.

We’ll also connect this to a problem many teams are already seeing: confidently incorrect answers in agent systems. When cache behavior, context reuse, and routing aren’t designed well, systems don’t just get slower or more expensive. They get inconsistent. And that shows up as wrong answers with high confidence. This session will walk through these trade-offs using live demos, showing how different KV cache strategies impact cost, latency, and output quality in real time.

Related talks at this congress

Open session

World Congress 2026 North America

September 24, 2026 · 16:10–16:40

Stage 6

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization

Legare Kerrison, Cedric Clyburn

Legare Kerrison
Cedric Clyburn
Open session

World Congress 2026 North America

September 24, 2026 · 14:10–14:40

Stage 1

Anatomy of an AI Request: Where Latency and Cost Are Really Born

Dan Fu

VP of Kernels at Together AI

Dan Fu
Open session

World Congress 2026 North America

September 24, 2026 · 10:20–10:50

Stage 5

Vector, Graph, or Key Value? Choosing Your Agent's Memory

Elizabeth Fuentes Leone

AWS - Developer Advocate/SDE, GenAI

Elizabeth Fuentes Leone
Open session

World Congress 2026 North America

September 25, 2026 · 13:30–14:00

Mainstage

A Hands-On Developer Guide to Inference Engineering

Ankit Patel, Philip Kiely

Ankit Patel
Philip Kiely
All sessions at this congress