World Congress 2026 North America • Sep 25, 2026 • Session details

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization

Legare Kerrison , Cedric Clyburn

Struggling to balance cost, accuracy, and latency in your LLM deployments? Discover how vLLM and quantization can slash your VRAM needs by 50% without sacrificing performance.

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization thumbnail

Checking access…

Playback and chapters load privately for Free videos.

Matching moments

2:39 min

Open source tools for running and scaling models

Cedric Clyburn Cedric Clyburn +1 · World Congress 2025

2:44 min

Running lightweight large language models on local hardware

Ekaterina Sirazitdinova · LIVE

2:00 min

Navigating the layers of the language model inference stack

Christin Pohl Christin Pohl · World Congress 2026 Europe

6:15 min

Executing open weight large language models with WebLLM

Christian Liebel Christian Liebel · World Congress 2025

1:14 min

Evaluating AI models using an LLM as a judge

Tomislav Tipurić Tomislav Tipurić +1 · World Congress 2026 North America

4:04 min

Evaluating model performance and accuracy using LLM judges

Viktoria Semaan Viktoria Semaan · World Congress 2026 North America