World Congress 2026 North America • Sep 25, 2026 • Session details

A Hands-On Developer Guide to Inference Engineering

Ankit Patel , Philip Kiely

Scaling web apps once meant serving kilobytes of HTML. Today, inference engineers orchestrate terabytes of model weights and gigabytes of KV cache to deploy high-throughput generative AI.

Pause
Mute Enter Fullscreen
#1 about 6 min

System-level perspective on generative artificial intelligence inference

Generative artificial intelligence services operate as complex systems involving harnesses, safety models, and routers rather than single models.

#2 about 5 min

Understanding pre-fill and decode phases in model latency

The transformer architecture splits computation into a compute-heavy intent phase and an autoregressive generation phase.

#3 about 3 min

Balancing latency and throughput in inference service engineering

Maximizing quality of service requires optimizing compute resources to serve fast answers without burning excessive capacity.

#4 about 3 min

Managing token generation efficiently using cache aware routing

Reusing cached input context across nodes significantly improves throughput and reduces costs for long-sequence agentic loops.

#5 about 3 min

Optimizing throughput and latency with model parallelism strategies

Choosing between tensor parallelism and expert parallelism balances the trade-offs between rapid generation and cost-effective scaling.

#6 about 2 min

Disaggregating pre-fill and decode workloads across compute clusters

Separating the initial prompt processing from autoregressive generation allows targeted parameter tuning across data center systems.

#7 about 2 min

Utilizing open source orchestration frameworks for cache management

Tools like NVIDIA Dynamo help manage resource allocation and double per-GPU throughput in disaggregated serving setups.

#8 about 6 min

Handling infrastructure scaling and cold start times efficiently

Rapidly pulling massive container images and terabytes of model weights prevents resource bottlenecks during unexpected traffic spikes.

#9 about 5 min

Building inference engineering intuition using accessible local hardware

Engineers can learn advanced optimization concepts by experimenting with small-parameter models on standard personal computers.

Matching moments

5:58 min

Architecting inference layers and meta harnesses to manage token costs

Hari Lingamagunta Hari Lingamagunta +3 · World Congress 2026 North America

50 sec

Key takeaways for optimizing open source inference engine deployments

Legare Kerrison Legare Kerrison +1 · World Congress 2026 North America

1:52 min

Prioritizing developer user experience over raw model parameters

Neel Sundaresan Neel Sundaresan +1 · World Congress 2026 Europe

2:54 min

Optimizing infrastructure for agentic flows and inference

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

1:20 min

Engineering lessons for scaling high-performance AI infrastructure

Wayne Liu Wayne Liu · World Congress 2026 North America

1:24 min

Exploring open source inference tools for language model serving

Legare Kerrison Legare Kerrison +1 · World Congress 2026 North America