World Congress 2026 North America
Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization
Legare Kerrison, Cedric Clyburn
World Congress 2026 North America
World Congress 2026 North America
September 23–25, 2026 · San José, CA
Attend in person
Get ticketsWatch remotely
Pro
Can’t make it to San José? Watch this session live with Pro. You also get:
Large Language Models are often described as if each generation introduces an entirely new architecture. In practice, most modern LLMs still retain the transformer core, but their real progress comes from a series of targeted design changes around attention, positional handling, feed-forward computation, routing, and memory efficiency. This talk explains LLM architectures through the engineering tradeoffs that shaped modern models: why some attention mechanisms evolved for lower inference cost, how architectural choices influence long-context behavior, why sparse activation changed the economics of scale, and how these shifts affect real-world deployment. Rather than treating LLM architecture as a static diagram, this session presents it as a set of design decisions made in response to practical constraints in latency, memory, context length, and system efficiency. Attendees will leave with a clearer mental model of what remained stable, what changed, and why those changes matter.
World Congress 2026 North America
Legare Kerrison, Cedric Clyburn
World Congress 2026 North America
Emmanuel Acheampong
Senior Manager Developer Relations at Crusoe AI
World Congress 2026 North America
An Phan
Senior Data Infrastructure Engineer @ Hippo Harvest
World Congress 2026 North America
Tejas Chopra
Senior Software Engineer at Netflix