> Markdown version of [/events/world-congress-2026-europe/sessions/1061-accelerating-ai](https://www.wearedevelopers.com/events/world-congress-2026-europe/sessions/1061-accelerating-ai). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Accelerating AI Inference at Scale: A Deep Dive Into NVIDIA Dynamo on Kubernetes - **Date:** Thursday, Jul 9, 2026 - **Time:** 13:00–15:00 (120 min) - **Room:** Room M2 (40 Seats) - **Event:** World Congress 2026 Europe ## Description As foundation models move toward deeper test-time computation, inference becomes the dominant scaling constraint. Latency, throughput, and cost are governed by a small set of forces: autoregressive decoding, KV-cache growth, memory bandwidth, and scheduling under contention. This workshop frames large-scale inference through these emerging laws of inference, starting from first principles and building toward real systems. Learners deploy NVIDIA Dynamo on Kubernetes to operate aggregated and disaggregated inference architectures using built-in KV-aware routing and scheduling. The outcome is a principled understanding of where inference time and money go — and how architectural choices bend those curves in production. Participants will deploy both aggregated and disaggregated inference on a 4xA100 node and compare the performance of the two. Workshop Preparation: - Please bring your own laptop. - Please review the following document and prepare accordingly before the workshop: <https://developer.nvidia.com/dli/getready> ## Speakers ### [Anshul Jindal](https://www.wearedevelopers.com/@anshul-jindal) Senior Solution Architect at NVIDIA ### [Mohak Chadha](https://www.wearedevelopers.com/@mohak-chadha) Solution Architect at NVIDIA ## Related talks at this congress - [Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs](https://www.wearedevelopers.com/events/world-congress-2026-europe/sessions/1350-agents-that-own) — Duan Lightfoot - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/events/world-congress-2026-europe/sessions/1387-running-secure-life) — Jeremy Murray - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/events/world-congress-2026-europe/sessions/1373-instant-kai) — Piotr Zaniewski - [Running AI at Scale: The Secret Ingredients](https://www.wearedevelopers.com/events/world-congress-2026-europe/sessions/1273-running-ai-at-scale) — Boris Hecker, Max Tschochohei, Peter Kürpick, Tomislav Tipurić