World Congress 2026 Europe - Virtual Stage • Jul 2, 2026 • Session details

The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)

Lerna Ekmekcioglu

Distributed AI training is a fragile seance where one dropped connection crashes the entire process. Discover how to eliminate network stragglers and build fault-tolerant GPU clusters at scale.

Pause
Mute Enter Fullscreen
#1 about 4 min

The AI workload technology stack and its components

The AI infrastructure stack relies on interconnected machine learning frameworks and network libraries to execute distributed operations.

#2 about 5 min

Anatomy of a distributed AI training loop in Kubernetes

Orchestrating tools execute forward passes, backward passes, and optimization steps across multiple containerized nodes.

#3 about 2 min

Strategies for scaling AI workloads across multiple GPUs

Data, tensor, and pipeline parallelism strategies divide compute tasks to optimize processing across massive hardware clusters.

#4 about 2 min

Coordinating GPU communication with NVIDIA collective communications library

Collective operations manage hardware topology and data synchronization to ensure model coherence across parallel operations.

#5 about 5 min

How slow queue pairs degrade AI training job throughput

High-latency network paths delay synchronization cycles and drastically reduce hardware utilization efficiency.

#6 about 4 min

Comparing backend and frontend networks for AI workloads

Dedicated architectures separate low-latency compute traffic from standard general-purpose management connectivity.

#7 about 2 min

Bypassing the CPU stack with remote direct memory access

Hardware-accelerated protocols transfer data directly between device memory blocks to eliminate kernel bottleneck latencies.

#8 about 5 min

Impact of network link flapping on distributed training jobs

Transient physical hardware disconnections immediately crash synchronous distributed machine learning sessions.

#9 about 4 min

Mitigating failure impact through regular AI model checkpointing

Persisting model state to storage at regular step boundaries prevents complete data loss during infrastructure outages.

Matching moments

3:05 min

Acquiring Mellanox to build cohesive AI factories

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

3:28 min

Challenges of shared Kubernetes clusters for AI workloads

Piotr Zaniewski Piotr Zaniewski · World Congress 2026 Europe

2:51 min

Designing hardware infrastructure and networking for distributed compute

Anshul Jindal Anshul Jindal +1 · World Congress 2025

4:26 min

Five unique failure modes in production AI systems

Heather Thacker Heather Thacker · Europe 2026 Virtual

1:38 min

Scaling bottlenecks in generative AI applications

Stan Girard Stan Girard · World Congress 2024

5:27 min

Simulating AI system cascades and graceful load shedding

Heather Thacker Heather Thacker · Europe 2026 Virtual