The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Distributed AI training is a fragile seance where one dropped connection crashes the entire process. Discover how to eliminate network stragglers and build fault-tolerant GPU clusters at scale.
Matching moments
More from World Congress 2026 Europe - Virtual Stage
Related videos
Related articles
BB
Benedikt Bischof
BB
Benedikt Bischof
DC
Daniel Cranney
MH
Michael Hunger
N
Neo4j
From learning to earning
Jobs that call for the skills explored in this talk.
about 1 month ago
•
Verified
GPU Cluster Engineer, Systems & Platform
Sciforium
San Francisco, United States
Expert
$150k–220k
Kubernetes
about 1 month ago
•
Verified
GPU Cluster Engineer, Networking
Sciforium
San Francisco, United States
Expert
$150k–180k
Network Infrastructure
about 1 month ago
•
Verified
Distributed Training and Inference Engineer
Sciforium
San Francisco, United States
Expert
$190k–250k
Linux kernel
16 days ago
Principal Software Engineer, AI Compute Infrastructure
ARM
Seattle, WA, United States
Expert
$262k
Linux
Grafana
Pytorch
about 1 month ago
•
Verified
LLM Training Engineer
Sciforium
San Francisco, United States
Expert
$155k–220k
Python
about 1 month ago
•
Verified
GPU Cluster Engineer, Hardware Operations
Sciforium
San Francisco, United States
Expert
$150k–180k
Network Infrastructure