Anatomy of an AI Request: Where Latency and Cost Are Really Born
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Running pre-fill and decode on the same GPU severely bottlenecks your AI throughput. Discover how cache-aware disaggregation and speculative decoding can drastically reduce latency and infrastructure costs.
Checking access…
Playback and chapters load privately for Free videos.
Matching moments
More from World Congress 2026 North America
Related videos
Related articles
DC
Daniel Cranney
BB
Benedikt Bischof
BB
Benedikt Bischof
MH
Michael Hunger
N
Neo4j
From learning to earning
Jobs that call for the skills explored in this talk.
26 days ago
Principal Software Engineer, AI Inference Runtime
ARM
Seattle, WA, United States
Expert
$262k
Compilers
Low Latency
Concurrency
about 2 months ago
•
Verified
LLM Training Engineer
Sciforium
San Francisco, United States
Expert
$155k–220k
Python
26 days ago
Staff Software Engineer, AI Inference Runtime
ARM
Seattle, WA, United States
Expert
$209k–282k
Compilers
Low Latency
Concurrency
26 days ago
Principal Software Engineer, AI Inference Cloud
ARM
Seattle, WA, United States
Expert
$262k
Pytorch
TensorRT
Kubernetes
about 2 months ago
•
Verified
GPU Kernel Engineer
Sciforium
San Francisco, United States
Expert
$190k–250k
PyTorch
about 2 months ago
•
Verified
Distributed Training and Inference Engineer
Sciforium
San Francisco, United States
Expert
$190k–250k
Linux kernel