World Congress 2026 North America
Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization
Legare Kerrison, Cedric Clyburn
Marek Suppa
Can you run massive NLP models within AWS Lambda's strict 250MB limit? Discover how Slido used knowledge distillation and ONNX to achieve sub-100ms serverless inference.
Jobs that call for the skills explored in this talk.