World Congress 2026 North America • Sep 24, 2026 • Session details

Anatomy of an AI Request: Where Latency and Cost Are Really Born

Dan Fu

Running pre-fill and decode on the same GPU severely bottlenecks your AI throughput. Discover how cache-aware disaggregation and speculative decoding can drastically reduce latency and infrastructure costs.

Anatomy of an AI Request: Where Latency and Cost Are Really Born thumbnail

Checking access…

Playback and chapters load privately for Free videos.

Matching moments

5:58 min

Architecting inference layers and meta harnesses to manage token costs

Hari Lingamagunta Hari Lingamagunta +3 · World Congress 2026 North America

1:47 min

Managing compute costs and AI model routing

Deivids Vilkinsons Deivids Vilkinsons +3 · World Congress 2026 Europe

3:20 min

Optimizing AI infrastructure from applications to GPU kernels

Erran Berger Erran Berger +1 · World Congress 2026 North America

1:52 min

Prioritizing developer user experience over raw model parameters

Neel Sundaresan Neel Sundaresan +1 · World Congress 2026 Europe

2:37 min

Understanding core parameters and mechanics of large language models

Julián Duque Julián Duque · World Congress 2025

1:47 min

Balancing scale and architecture in artificial intelligence development

Yuval Dvir Yuval Dvir · World Congress 2026 North America