World Congress 2026 North America
Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization
Legare Kerrison, Cedric Clyburn
World Congress 2026 North America
World Congress 2026 North America
September 23–25, 2026 · San José, CA
Attend in person
Get ticketsWatch remotely
Pro
Can’t make it to San José? Watch this session live with Pro. You also get:
Most teams think KV cache is about making inference faster. That’s true, but it’s not the point. KV cache is really about controlling memory, reducing recomputation, and keeping costs from spiraling as usage grows. As models get deployed at scale, the real bottleneck is no longer raw compute. It’s memory, bandwidth, and power. This session takes a step back to explain what KV cache does at a system level, why it matters for real workloads, and how approaches like vLLM, LMCache, and SGLang change how we think about scaling inference.
We’ll also connect this to a problem many teams are already seeing: confidently incorrect answers in agent systems. When cache behavior, context reuse, and routing aren’t designed well, systems don’t just get slower or more expensive. They get inconsistent. And that shows up as wrong answers with high confidence. This session will walk through these trade-offs using live demos, showing how different KV cache strategies impact cost, latency, and output quality in real time.
World Congress 2026 North America
Legare Kerrison, Cedric Clyburn
World Congress 2026 North America
Duan Lightfoot
Sr. AI Engineer, Akamai
World Congress 2026 North America
Carl Lapierre
Tech Lead and AI Engineer at Osedea
World Congress 2026 North America
Jofia Jose Prakash
Enterprise AI Architect at American Chemical Society