> Markdown version of [/videos/100542-a-hands-on-developer-guide-to-inference-engineering](https://www.wearedevelopers.com/videos/100542-a-hands-on-developer-guide-to-inference-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # A Hands-On Developer Guide to Inference Engineering Scaling web apps once meant serving kilobytes of HTML. Today, inference engineers orchestrate terabytes of model weights and gigabytes of KV cache to deploy high-throughput generative AI. - **Speakers:** [Ankit Patel](https://www.wearedevelopers.com/@ankit-patel), [Philip Kiely](https://www.wearedevelopers.com/@philip-kiely) - **Event:** World Congress 2026 North America - **Published:** September 25, 2026 - **Duration:** 30:18 - **URL:** https://www.wearedevelopers.com/videos/100542-a-hands-on-developer-guide-to-inference-engineering ## Summary Deploying generative AI is rarely as simple as interacting with a single frontier model; it requires navigating a complex "systems problem" involving safety harnesses, routers, and multiple LLMs conversing via prompting. Inference engineers must solve a multi-dimensional puzzle balancing latency, throughput, and operational cost. A foundational challenge lies in the transformer architecture itself, which demands compute-heavy "pre-fill" stages to process intent, followed by autoregressive "decode" phases to generate output tokens. Balancing the time to first token against high-throughput capacity forces teams to finely tune the serving stack according to specific application demands. To scale effectively, especially for agentic workflows with massive context windows, engineers must master the "KV cache hot potato"—efficiently storing and transferring cached input tokens. Implementing KV-aware routing ensures requests are directed to replicas holding the necessary cache, drastically improving hit rates and reducing compute overhead. Furthermore, operators must select optimal parallelism strategies for trillion-parameter, mixture of experts models. Tensor parallelism heavily optimizes latency for speed-critical applications, whereas expert parallelism maximizes throughput for broader, cost-effective service delivery. Advanced enterprise setups utilize prefill-decode disaggregation to assign dedicated hardware to each phase, a technique shown to double overall token throughput when managed by open-source orchestration frameworks like NVIDIA Dynamo. Maintaining these sophisticated systems in production introduces massive infrastructure hurdles. The classical web engineering challenge of scaling to serve kilobytes of HTML has evolved into orchestrating "terabytes of model weights and gigabytes of KV cache." Mitigating the thundering herd problem during sudden traffic spikes requires drastically reducing cold start times via high-bandwidth pipes and standardized global compute pools. Ultimately, as reinforcement learning embeds high-throughput inference directly into model training, the discipline of inference engineering becomes critical across the entire AI lifecycle. Practitioners can build essential intuition by experimenting locally with tools like vLLM, TensorRT-LLM, and SGLang before scaling up to multi-cluster production environments. **Keywords:** inference engineering, generative AI serving, prefill-decode disaggregation, KV cache management, KV-aware routing, time to first token, tensor parallelism, expert parallelism, mixture of experts, GPU cold start optimization, autoregressive decoding, nvidia dynamo orchestration, vLLM framework, tensorrt-LLM, LLM serving infrastructure ## Chapters 1. **System-level perspective on generative artificial intelligence inference** (00:00) — Generative artificial intelligence services operate as complex systems involving harnesses, safety models, and routers rather than single models. 1. **Understanding pre-fill and decode phases in model latency** (05:13) — The transformer architecture splits computation into a compute-heavy intent phase and an autoregressive generation phase. 1. **Balancing latency and throughput in inference service engineering** (09:21) — Maximizing quality of service requires optimizing compute resources to serve fast answers without burning excessive capacity. 1. **Managing token generation efficiently using cache aware routing** (11:38) — Reusing cached input context across nodes significantly improves throughput and reduces costs for long-sequence agentic loops. 1. **Optimizing throughput and latency with model parallelism strategies** (14:13) — Choosing between tensor parallelism and expert parallelism balances the trade-offs between rapid generation and cost-effective scaling. 1. **Disaggregating pre-fill and decode workloads across compute clusters** (17:04) — Separating the initial prompt processing from autoregressive generation allows targeted parameter tuning across data center systems. 1. **Utilizing open source orchestration frameworks for cache management** (18:48) — Tools like NVIDIA Dynamo help manage resource allocation and double per-GPU throughput in disaggregated serving setups. 1. **Handling infrastructure scaling and cold start times efficiently** (20:38) — Rapidly pulling massive container images and terabytes of model weights prevents resource bottlenecks during unexpected traffic spikes. 1. **Building inference engineering intuition using accessible local hardware** (25:44) — Engineers can learn advanced optimization concepts by experimenting with small-parameter models on standard personal computers. ## Related Moments - [Architecting inference layers and meta harnesses to manage token costs](https://www.wearedevelopers.com/videos/100600-inside-the-ai-native-engineering-org) (from "Inside the AI-Native Engineering Org") - [Key takeaways for optimizing open source inference engine deployments](https://www.wearedevelopers.com/videos/100416-fast-cheap-and-accurate-optimizing-llm-inference-with-vllm-and-quantization) (from "Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization") - [Prioritizing developer user experience over raw model parameters](https://www.wearedevelopers.com/videos/100256-can-this-elephant-dance-ibm-bob-and-the-future-of-ai-first-software-development) (from "Can This Elephant Dance? IBM Bob and the Future of AI-First Software Development") - [Optimizing infrastructure for agentic flows and inference](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) (from "Building the Nervous System of AI - Michael Kagan (NVIDIA)") - [Engineering lessons for scaling high-performance AI infrastructure](https://www.wearedevelopers.com/videos/100365-designing-high-performance-ai-apis-lessons-from-serving-millions-of-real-time-requests) (from "Designing High-Performance AI APIs: Lessons from Serving Millions of Real-Time Requests") - [Exploring open source inference tools for language model serving](https://www.wearedevelopers.com/videos/100416-fast-cheap-and-accurate-optimizing-llm-inference-with-vllm-and-quantization) (from "Fast, Cheap, and Accurate: Optimizing LLM Inference with vLLM and Quantization") ## Related Articles - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) ## Related Jobs - [Principal Software Engineer, AI Inference Cloud](https://www.wearedevelopers.com/jobs/ext/2854957-principal-software-engineer-ai-inference-cloud) at **ARM** - [Staff Software Engineer, AI Inference Cloud](https://www.wearedevelopers.com/jobs/ext/3347267-staff-software-engineer-ai-inference-cloud) at **ARM** - [Principal Software Engineer, AI Inference Runtime](https://www.wearedevelopers.com/jobs/ext/2854958-principal-software-engineer-ai-inference-runtime) at **ARM** - [Staff Software Engineer, AI Inference Runtime](https://www.wearedevelopers.com/jobs/ext/3474466-staff-software-engineer-ai-inference-runtime) at **ARM** - [LLM Training Engineer](https://www.wearedevelopers.com/jobs/48420-llm-training-engineer) at **Sciforium** - [Distributed Training and Inference Engineer](https://www.wearedevelopers.com/jobs/48399-distributed-training-and-inference-engineer) at **Sciforium**