> Markdown version of [/videos/100416-fast-cheap-and-accurate-optimizing-llm-inference-with-vllm-and-quantization?t=1536](https://www.wearedevelopers.com/videos/100416-fast-cheap-and-accurate-optimizing-llm-inference-with-vllm-and-quantization?t=1536). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization Struggling to balance cost, accuracy, and latency in your LLM deployments? Discover how vLLM and quantization can slash your VRAM needs by 50% without sacrificing performance. - **Speakers:** [Legare Kerrison](https://www.wearedevelopers.com/@technicallylegare), [Cedric Clyburn](https://www.wearedevelopers.com/@cedric-clyburn) - **Event:** World Congress 2026 North America - **Published:** September 25, 2026 - **Duration:** 26:27 - **URL:** https://www.wearedevelopers.com/videos/100416-fast-cheap-and-accurate-optimizing-llm-inference-with-vllm-and-quantization ## Access Playback and chapters for this video are available with a Free account. ## Related Moments - [Open source tools for running and scaling models](https://www.wearedevelopers.com/videos/1597-self-hosted-llms-from-zero-to-inference) (from "Self-Hosted LLMs: From Zero to Inference") - [Running lightweight large language models on local hardware](https://www.wearedevelopers.com/videos/829-multimodal-generative-ai-demystified) (from "Multimodal Generative AI Demystified") - [Navigating the layers of the language model inference stack](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) (from "Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated") - [Executing open weight large language models with WebLLM](https://www.wearedevelopers.com/videos/1615-prompt-api-webnn-the-ai-revolution-right-in-your-browser) (from "Prompt API & WebNN: The AI Revolution Right in Your Browser") - [Evaluating AI models using an LLM as a judge](https://www.wearedevelopers.com/videos/100508-legacy-as-a-launchpad-how-yahoo-mail-is-undergoing-a-product-and-engineering-transformation) (from "Legacy as a Launchpad: How Yahoo Mail is Undergoing a Product and Engineering Transformation") - [Evaluating model performance and accuracy using LLM judges](https://www.wearedevelopers.com/videos/100447-from-model-selection-to-smart-routing-how-to-use-the-right-llm-for-every-task) (from "From Model Selection to Smart Routing: How to Use the Right LLM for Every Task") ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) ## Related Jobs - [LLM Training Engineer](https://www.wearedevelopers.com/jobs/48420-llm-training-engineer) at **Sciforium** - [LLM Dataset Engineer](https://www.wearedevelopers.com/jobs/48419-llm-dataset-engineer) at **Sciforium** - [Model Implementation Engineer](https://www.wearedevelopers.com/jobs/48421-model-implementation-engineer) at **Sciforium** - [Lead Software Engineer, Model Serving Platform](https://www.wearedevelopers.com/jobs/48413-lead-software-engineer-model-serving-platform) at **Sciforium** - [ML Engineer](https://www.wearedevelopers.com/jobs/48422-ml-engineer) at **Sciforium** - [ML Engineer](https://www.wearedevelopers.com/jobs/48448-ml-engineer) at **Docker, Inc.**