World Congress 2026 North America
Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs
Duan Lightfoot
Sr. AI Engineer, Akamai
World Congress 2026 North America
World Congress 2026 North America
September 23–25, 2026 · San José, CA
Attend in person
Get ticketsWatch remotely
Pro
Can’t make it to San José? Watch this session live with Pro. You also get:
Inference engineers solve a multi-dimensional puzzle across latency, throughput, and cost to serve generative AI models. Spanning interdependent layers of the serving stack, from CUDA to runtimes to containers to Kubernetes, inference engineering is the discipline behind scaling AI applications. In this session, we’ll cover key inference engineering concepts across both runtime and infrastructure, including topology-aware model parallelism, prefill-decode disaggregated serving, KV-aware routing, autoscaling strategies, and multi-cluster infrastructure management, with a focus on the work required to run multi-trillion-parameter LLMs efficiently in production.
World Congress 2026 North America
Duan Lightfoot
Sr. AI Engineer, Akamai
World Congress 2026 North America
Kyle Bell
VP of AI @ TensorWave
World Congress 2026 North America
An Phan
Senior Data Infrastructure Engineer @ Hippo Harvest
World Congress 2026 North America
David vonThenen
AI/ML Leader | Keynote Speaker | OSS Engineer & Developer Advocate | Agentic AI, Deep Learning, Production AI | Python, Go, C++