World Congress 2024 Aug 20, 2024 Session details

AI Factories at Scale

Thomas Schmidt

Generative AI demands a shift from legacy CPUs to high-density GPU clusters. Learn how to architect scalable AI factories that slash compute energy usage.

Pause
Mute Enter Fullscreen
#1 about 3 min

Amber's evolution as a pioneer in GPU acceleration

How early fluid dynamics research evolved into a foundational enterprise partnership with Nvidia for AI infrastructure.

#2 about 3 min

Milestones driving the rapid evolution of generative AI

Landmark computing breakthroughs like CUDA and transformers dramatically accelerated the foundational capabilities of artificial intelligence models.

#3 about 3 min

Financial impacts of generative AI across the enterprise

Pushing early generative artificial intelligence experimentation into real-world use cases creates significant return on investment for enterprises.

#4 about 2 min

Contrasting legacy supercomputers with modern AI clustered infrastructure

The modern DGX H100 provides exponential performance improvements at a fraction of the cost and compute footprint of traditional supercomputers.

#5 about 3 min

Replacing legacy CPUs to reduce data center energy consumption

Transitioning high-intensity workloads from CPUs to specialized GPUs significantly lowers global energy constraints while preventing thermal overload.

#6 about 2 min

Escalating compute demands for production generative AI inference

The transition toward production software integration multiplies infrastructural demands required to serve transformer models asynchronously.

#7 about 3 min

Core infrastructure components required for an AI factory

Essential deployment layers combine internal cluster networking, advanced storage structures, and modular temperature controls to finalize AI architecture.

#8 about 1 min

Deploying an immediate AI center of excellence via superpods

Turnkey clusters establish scalable reference structures directly optimized for massively parallel application training deployments.

#9 about 5 min

Managing the AI cluster using stack synchronization software

Dedicated management software resolves complex stack synchronization, coordinates hybrid cloud workloads, and proactively monitors cluster networking health.

#10 about 1 min

Allocating GPU resources and implementing multi-tenant usage chargebacks

Defining specific multi-instance constraints permits scalable resource allocation while allowing enterprises to accurately attribute processing costs to specific users.

#11 about 2 min

High-performance storage necessities for continuous large LLM training

Specialized high-bandwidth storage tiers bypass input bottlenecks that frequently stall training sequences during comprehensive model checkpointing.

#12 about 3 min

Adapting data center environments for direct liquid immersion cooling

Upgrading physical infrastructure arrays toward integrated immersion tanks safely manages the vast thermal footprints generated by emerging inference parameters.

Matching moments

1:25 min

Addressing the sustainability and power consumption of AI

Christian Heilmann Christian Heilmann · WWC Europe 2026

2:12 min

Transitioning from data centers to AI factories

Michael Kagan Michael Kagan +1 · WWC Europe 2026

4:03 min

Managing massive power consumption scaling in AI data centers

Stephan Gillich Stephan Gillich +3 · WWC 2024

2:35 min

Balancing AI competitiveness with compute efficiency demands

Markus Hacker Markus Hacker +1 · WWC 2025

3:57 min

Strategies for accelerating innovation and maximizing AI value

Stephan Gillich Stephan Gillich · WWC 2024

1:38 min

Scaling bottlenecks in generative AI applications

Stan Girard Stan Girard · WWC 2024

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

Autonomous Infrastructure: Building AI Agents for Global-Scale Capacity Efficiency

Gregoire Colin, Tommy Tran

Gregoire Colin
Tommy Tran
Open session

World Congress 2026 North America

Building Stuff with GenAI - The Open Minded Workshop beyond OpenAI

Andreas Erben

CTO for Applied AI and Metaverse at daenet

Andreas Erben
Open session

World Congress 2026 North America

Trust, But Verify: Continuous GPU Validation at Scale

Kyle Bell

VP of AI @ TensorWave

Kyle Bell
Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

Compute for your AI model: GPUs, LPUs, TPUs and beyond..

Kushaagra Goyal

Tech Lead at Rubrik, ex-CTO at Gan.AI, ex-Databricks

Kushaagra Goyal