World Congress 2026 Europe - Virtual Stage Jul 2, 2026 Session details

The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)

Lerna Ekmekcioglu

Distributed AI training is a fragile seance where one dropped connection crashes the entire process. Discover how to eliminate network stragglers and build fault-tolerant GPU clusters at scale.

Pause
Mute Enter Fullscreen
#1 about 4 min

The AI workload technology stack and its components

The AI infrastructure stack relies on interconnected machine learning frameworks and network libraries to execute distributed operations.

#2 about 5 min

Anatomy of a distributed AI training loop in Kubernetes

Orchestrating tools execute forward passes, backward passes, and optimization steps across multiple containerized nodes.

#3 about 2 min

Strategies for scaling AI workloads across multiple GPUs

Data, tensor, and pipeline parallelism strategies divide compute tasks to optimize processing across massive hardware clusters.

#4 about 2 min

Coordinating GPU communication with NVIDIA collective communications library

Collective operations manage hardware topology and data synchronization to ensure model coherence across parallel operations.

#5 about 5 min

How slow queue pairs degrade AI training job throughput

High-latency network paths delay synchronization cycles and drastically reduce hardware utilization efficiency.

#6 about 4 min

Comparing backend and frontend networks for AI workloads

Dedicated architectures separate low-latency compute traffic from standard general-purpose management connectivity.

#7 about 2 min

Bypassing the CPU stack with remote direct memory access

Hardware-accelerated protocols transfer data directly between device memory blocks to eliminate kernel bottleneck latencies.

#8 about 5 min

Impact of network link flapping on distributed training jobs

Transient physical hardware disconnections immediately crash synchronous distributed machine learning sessions.

#9 about 4 min

Mitigating failure impact through regular AI model checkpointing

Persisting model state to storage at regular step boundaries prevents complete data loss during infrastructure outages.

Matching moments

3:05 min

Acquiring Mellanox to build cohesive AI factories

Michael Kagan Michael Kagan +1 · WWC Europe 2026

3:28 min

Challenges of shared Kubernetes clusters for AI workloads

Piotr Zaniewski Piotr Zaniewski · WWC Europe 2026

2:51 min

Designing hardware infrastructure and networking for distributed compute

Anshul Jindal Anshul Jindal +1 · WWC 2025

4:26 min

Five unique failure modes in production AI systems

Heather Thacker Heather Thacker · Europe 2026 Virtual

1:38 min

Scaling bottlenecks in generative AI applications

Stan Girard Stan Girard · WWC 2024

5:27 min

Simulating AI system cascades and graceful load shedding

Heather Thacker Heather Thacker · Europe 2026 Virtual

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Trust, But Verify: Continuous GPU Validation at Scale

Kyle Bell

VP of AI @ TensorWave

Kyle Bell
Open session

World Congress 2026 North America

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

No Single Model to Rule Them All: Building Resilient AI Agents Across Open & Closed LLMs

Emmanuel Acheampong

Senior Manager Developer Relations at Crusoe AI

Emmanuel Acheampong
Open session

World Congress 2026 North America

You Can’t Re-Run Sunlight: Designing ML Data Architectures for Physical AI

An Phan

Senior Data Infrastructure Engineer @ Hippo Harvest

An Phan
Open session

World Congress 2026 North America

Compute for your AI model: GPUs, LPUs, TPUs and beyond..

Kushaagra Goyal

Tech Lead at Rubrik, ex-CTO at Gan.AI, ex-Databricks

Kushaagra Goyal
Open session

World Congress 2026 North America

Autonomous Infrastructure: Building AI Agents for Global-Scale Capacity Efficiency

Tommy Tran, Gregoire Colin

Tommy Tran
Gregoire Colin