> Markdown version of [/videos/100526-anatomy-of-an-ai-request-where-latency-and-cost-are-really-born](https://www.wearedevelopers.com/videos/100526-anatomy-of-an-ai-request-where-latency-and-cost-are-really-born). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Anatomy of an AI Request: Where Latency and Cost Are Really Born Running pre-fill and decode on the same GPU severely bottlenecks your AI throughput. Discover how cache-aware disaggregation and speculative decoding can drastically reduce latency and infrastructure costs. - **Speakers:** [Dan Fu](https://www.wearedevelopers.com/@dan-fu) - **Event:** World Congress 2026 North America - **Published:** September 24, 2026 - **Duration:** 28:01 - **URL:** https://www.wearedevelopers.com/videos/100526-anatomy-of-an-ai-request-where-latency-and-cost-are-really-born ## Access Playback and chapters for this video are available with a Free account. ## Related Moments - [Architecting inference layers and meta harnesses to manage token costs](https://www.wearedevelopers.com/videos/100600-inside-the-ai-native-engineering-org) (from "Inside the AI-Native Engineering Org") - [Managing compute costs and AI model routing](https://www.wearedevelopers.com/videos/100328-the-limits-of-llms-in-real-world-applications) (from "The Limits of LLMs in Real-World Applications") - [Optimizing AI infrastructure from applications to GPU kernels](https://www.wearedevelopers.com/videos/100536-how-linkedin-turns-ai-breakthroughs-into-member-and-customer-value) (from "How LinkedIn Turns AI Breakthroughs into member and customer value") - [Prioritizing developer user experience over raw model parameters](https://www.wearedevelopers.com/videos/100256-can-this-elephant-dance-ibm-bob-and-the-future-of-ai-first-software-development) (from "Can This Elephant Dance? IBM Bob and the Future of AI-First Software Development") - [Understanding core parameters and mechanics of large language models](https://www.wearedevelopers.com/videos/1512-building-ai-applications-with-langchain-and-node-js) (from "Building AI Applications with LangChain and Node.js") - [Balancing scale and architecture in artificial intelligence development](https://www.wearedevelopers.com/videos/100476-making-science-larger-not-just-faster) (from "Making Science Larger, not just Faster") ## Related Articles - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) ## Related Jobs - [Principal Software Engineer, AI Inference Runtime](https://www.wearedevelopers.com/jobs/ext/2854958-principal-software-engineer-ai-inference-runtime) at **ARM** - [LLM Training Engineer](https://www.wearedevelopers.com/jobs/48420-llm-training-engineer) at **Sciforium** - [Staff Software Engineer, AI Inference Runtime](https://www.wearedevelopers.com/jobs/ext/3474466-staff-software-engineer-ai-inference-runtime) at **ARM** - [Principal Software Engineer, AI Inference Cloud](https://www.wearedevelopers.com/jobs/ext/2854957-principal-software-engineer-ai-inference-cloud) at **ARM** - [GPU Kernel Engineer](https://www.wearedevelopers.com/jobs/48412-gpu-kernel-engineer) at **Sciforium** - [Distributed Training and Inference Engineer](https://www.wearedevelopers.com/jobs/48399-distributed-training-and-inference-engineer) at **Sciforium**