World Congress 2026 North America • Sep 25, 2026 • Session details

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization

Legare Kerrison , Cedric Clyburn

Struggling to balance cost, accuracy, and latency in your LLM deployments? Discover how vLLM and quantization can slash your VRAM needs by 50% without sacrificing performance.

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization thumbnail

Checking access…

Playback and chapters load privately for Free videos.

Matching moments

2:39 min

Open source tools for running and scaling models

Cedric Clyburn Cedric Clyburn +1 · World Congress 2025

2:44 min

Running lightweight large language models on local hardware

Ekaterina Sirazitdinova · LIVE

2:00 min

Navigating the layers of the language model inference stack

Christin Pohl Christin Pohl · World Congress 2026 Europe

6:15 min

Executing open weight large language models with WebLLM

Christian Liebel Christian Liebel · World Congress 2025

1:14 min

Evaluating AI models using an LLM as a judge

Tomislav Tipurić Tomislav Tipurić +1 · World Congress 2026 North America

4:04 min

Evaluating model performance and accuracy using LLM judges

Viktoria Semaan Viktoria Semaan · World Congress 2026 North America

Upcoming sessions on this topic

Open session

Build a Zero-Cost AI Agent in Your Browser

  • Aishwarya Mathuria

    Adobe

    Computer Scientist

  • Nishant Thakur

    Adobe

    Senior Computer Scientist

Open session

Architecting Trusted AI: Coupling Multi-Agent Systems with dbt Semantic Store for Zero-Hallucination

  • Pushkar Mishra

    J.P. Morgan

    Senior Vice President

Open session

AI Is the Easy Part: Why Faster Engineers Don't Make Faster Businesses

  • Bharat Sharma

    Booking

    Senior Engineering Manager

Open session

Beyond File Dumps: Context Engineering for Coding Agents

  • Animesh Dutta

    Arm

    Senior Software Engineer

Open session

Java is Coming for Python’s Lunch: The Rise of JVM AI Agents

  • Frederik Pietzko

    Jetbrains

    Developer Advocate

Open session

What I Got Wrong Shipping an MCP Server for Live Infrastructure

  • Pritesh Kiri

    Harness

    Developer Relations Engineer