World Congress 2026 North America
Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization
Legare Kerrison, Cedric Clyburn
Stop overpaying for massive, generic LLMs. Standard engineering teams can use InstructLab to locally fine-tune open-source models, drastically cutting inference costs without specialized data science expertise.
Jobs that call for the skills explored in this talk.