World Congress 2026 North America
Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization
Legare Kerrison, Cedric Clyburn
Jodie Burchell
Are nested loops turning your Python scripts into multi-hour bottlenecks? Compress execution times down to mere milliseconds using NumPy and linear algebra vectorization.
Jobs that call for the skills explored in this talk.