World Congress 2026 North America
September 24, 2026 · 14:10–14:40
Stage 1
Anatomy of an AI Request: Where Latency and Cost Are Really Born
Dan Fu
VP of Kernels at Together AI
Dave Anderson , Sarel Weinberger , Phd
Compressing LLM inputs doesn't save money. It forces models to work harder, driving up compute costs by 50%. Learn smarter resource allocation techniques that actually lower your bills.
Jobs that call for the skills explored in this talk.