World Congress 2026 North America

Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization

September 23–25, 2026

World Congress 2026 North America

September 23–25, 2026 · San José, CA

Attend in person

Get tickets

Watch remotely

Watch live with Pro

Pro

Can’t make it to San José? Watch this session live with Pro. You also get:

  • All full videos, bookmarks, and playlists
  • World Congress livestreams
See pricing

What this session covers

Running large language models at scale can get expensive fast, but the right optimizations can cut latency and GPU costs dramatically.

We’ll walk through how to serve models efficiently using vLLM, an open-source, high-performance inference engine. Then and generate and test quantized models, expose them through vLLM’s OpenAI-compatible API, and tune runtime flags to balance throughput, latency, and accuracy on different GPUs.

We’ll benchmark performance live, inspect token-throughput metrics, and discuss real-world deployment trade-offs.

Related talks at this congress

Open session

World Congress 2026 North America

Understanding LLM Architectures: Inside the Design of Modern Models

Jofia Jose Prakash

Enterprise AI Architect at American Chemical Society

Jofia Jose Prakash
Open session

World Congress 2026 North America

KV Cache Is Not About Speed: It's About Surviving Inference Costs

David vonThenen

AI/ML Leader | Keynote Speaker | OSS Engineer & Developer Advocate | Agentic AI, Deep Learning, Production AI | Python, Go, C++

David vonThenen
Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
All sessions at this congress