World Congress 2026 North America • Sep 24, 2026 • Session details

Compute for your AI model: GPUs, LPUs, TPUs and beyond..

Kushaagra Goyal

Is your shared hardware starving your AI model's decode phase? Learn to slash inference latency using disaggregated serving across specialized GPUs, TPUs, and SRAM-dense LPUs.

Compute for your AI model: GPUs, LPUs, TPUs and beyond.. thumbnail

Checking access…

Playback and chapters load privately for Free videos.

Matching moments

1:47 min

Balancing scale and architecture in artificial intelligence development

Yuval Dvir Yuval Dvir · World Congress 2026 North America

3:18 min

Hardware architectures tailored for specific artificial intelligence computations

Stephan Gillich Stephan Gillich · World Congress 2024

1:32 min

Escalating compute demands for production generative AI inference

Thomas Schmidt Thomas Schmidt · World Congress 2024

2:07 min

Lowering inference costs through specialized hardware architectures

Ed Huang Ed Huang +4 · World Congress 2026 North America

2:00 min

Selecting purpose-built hardware for model training and inference

Sohan Maheshwar · LIVE

1:47 min

Managing compute costs and AI model routing

Deivids Vilkinsons Deivids Vilkinsons +3 · World Congress 2026 Europe