World Congress 2026 North America
September 24, 2026 · 16:10–16:40
Stage 9
Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization
Legare Kerrison, Cedric Clyburn
Ditch the 20-minute cold starts. Treat self-hosted LLMs as compute-heavy REST APIs. Learn how NFS, Pingora, and intelligent rate-limiting seamlessly scale your infrastructure to serve millions.
Jobs that call for the skills explored in this talk.