World Congress 2026 North America
Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization
Legare Kerrison, Cedric Clyburn
Maximizing open-source LLM throughput requires aggressively balancing compute and memory bottlenecks. Master the complete GPU optimization stack, from simple model quantization to sophisticated speculative decoding.
Jobs that call for the skills explored in this talk.