> Markdown version of [/events/world-congress-2026-north-america/sessions/1835-anatomy-of-an-ai](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1835-anatomy-of-an-ai). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Anatomy of an AI Request: Where Latency and Cost Are Really Born - **Date:** Thursday, Sep 24, 2026 - **Time:** 14:10–14:40 (30 min) - **Room:** Stage 1 - **Event:** World Congress 2026 North America ## Description Every LLM API call looks simple on the surface. But under the hood, it’s a deeply layered systems problem spanning kernels, compilers, GPU scheduling, and distributed inference infrastructure. In this session, Dan Fu, VP of Kernels at Together AI, breaks down what actually happens when a request hits a modern AI model—and how inference performance is ultimately determined. He walks through the full inference stack, from tokenization and model execution to GPU kernel dispatch, memory movement, and serving-time orchestration, highlighting where inefficiencies accumulate and why today’s systems operate far below theoretical hardware capability. Drawing on both cutting-edge systems research and production-scale infrastructure experience—including foundational work like FlashAttention, now widely adopted across the AI ecosystem—Dan unpacks the core bottlenecks in modern inference stacks. He also outlines where the largest efficiency gains are still available, and why kernel-level optimization is becoming a critical lever in scaling AI systems. ## Speaker ### [Dan Fu](https://www.wearedevelopers.com/@dan-fu) VP of Kernels at Together AI ## Related talks at this congress - [Designing High-Performance AI APIs: Lessons from Serving Millions of Real-Time Requests](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1846-designing-high) — Wayne Liu - [Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1717-agents-that-own) — Duan Lightfoot - [You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1677-you-can-t-re-run) — An Phan - [A Hands-On Developer Guide to Inference Engineering](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1771-a-hands-on-developer) — Ankit Patel, Philip Kiely ## Watch remotely Can’t make it to San José? Watch this session live with Pro. You also get: - All full videos, bookmarks, and playlists - World Congress livestreams [See pricing](https://www.wearedevelopers.com/pricing) ## Links - [Get tickets](https://www.wearedevelopers.com/world-congress-north-america/tickets)