World Congress 2026 North America

A Hands-On Developer Guide to Inference Engineering

September 23–25, 2026

World Congress 2026 North America

September 23–25, 2026 · San José, CA

Attend in person

Get tickets

Watch remotely

Watch live with Pro

Pro

Can’t make it to San José? Watch this session live with Pro. You also get:

  • All full videos, bookmarks, and playlists
  • World Congress livestreams
See pricing

What this session covers

Inference engineers solve a multi-dimensional puzzle across latency, throughput, and cost to serve generative AI models. Spanning interdependent layers of the serving stack, from CUDA to runtimes to containers to Kubernetes, inference engineering is the discipline behind scaling AI applications. In this session, we’ll cover key inference engineering concepts across both runtime and infrastructure, including topology-aware model parallelism, prefill-decode disaggregated serving, KV-aware routing, autoscaling strategies, and multi-cluster infrastructure management, with a focus on the work required to run multi-trillion-parameter LLMs efficiently in production.

Related talks at this congress

Open session

World Congress 2026 North America

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

Trust, But Verify: Continuous GPU Validation at Scale

Kyle Bell

VP of AI @ TensorWave

Kyle Bell
Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

KV Cache Is Not About Speed: It's About Surviving Inference Costs

David vonThenen

AI/ML Leader | Keynote Speaker | OSS Engineer & Developer Advocate | Agentic AI, Deep Learning, Production AI | Python, Go, C++

David vonThenen
All sessions at this congress