World Congress 2026 North America
Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization
Legare Kerrison, Cedric Clyburn
Why risk data privacy with third-party APIs when you can run models natively? Discover how to deploy, quantize, and scale self-hosted LLMs for completely secure, offline development.
Jobs that call for the skills explored in this talk.