World Congress 2026 North America • Sep 25, 2026 • Session details

The Unit Economics of AI

Ed Huang , Jeremy Murray , John Malcolm , Aaron Bawcom , Matt Burns

Are skyrocketing GPU costs stalling your AI deployments? Discover how intelligent routing, agentic caching, and model compression can slash inference bills and reduce your compute footprint by 80 percent.

The Unit Economics of AI thumbnail

Checking access…

Playback and chapters load privately for Free videos.

Matching moments

5:58 min

Architecting inference layers and meta harnesses to manage token costs

Hari Lingamagunta Hari Lingamagunta +3 · World Congress 2026 North America

2:26 min

Cost and latency pressures pushing AI to the edge

Moe Sani Moe Sani · World Congress 2026 Europe

3:21 min

Reducing cloud dependency with on-device edge AI models

Precious Osaro Precious Osaro · World Congress 2026 Europe

3:07 min

Managing economic costs and intelligence scaling in AI products

Logan Kilpatrick · Coffee With Developers

1:47 min

Managing compute costs and AI model routing

Deivids Vilkinsons Deivids Vilkinsons +3 · World Congress 2026 Europe

5:27 min

Externalized computing costs of AI scaling

Perf + AI