Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Struggling to balance cost, accuracy, and latency in your LLM deployments? Discover how vLLM and quantization can slash your VRAM needs by 50% without sacrificing performance.
Checking access…
Playback and chapters load privately for Free videos.
Matching moments
More from World Congress 2026 North America
Related videos
From learning to earning
Jobs that call for the skills explored in this talk.
about 2 months ago
•
Verified
LLM Training Engineer
Sciforium
San Francisco, United States
Expert
$155k–220k
Python
about 2 months ago
•
Verified
LLM Dataset Engineer
Sciforium
San Francisco, United States
Expert
$155k–210k
Python
about 2 months ago
•
Verified
Model Implementation Engineer
Sciforium
San Francisco, United States
Expert
$165k–220k
Python
about 2 months ago
•
Verified
Lead Software Engineer, Model Serving Platform
Sciforium
San Francisco, United States
Expert
$230k–300k
Python
about 2 months ago
•
Verified
ML Engineer
Sciforium
San Francisco, United States
Expert
$165k–210k
Python
about 1 month ago
•
Verified
ML Engineer
Docker, Inc.
Seattle, United States
Expert
Remote
Go