World Congress 2025
July 10, 2025 · 13:30–14:00
Stage 5
Your Next AI Needs 10,000 GPUs. Now What?
Anshul Jindal, Martin Piercy
World Congress 2025
Kubernetes traditionally does not have a mechanism for allocating non-node-local resources. The one exception being persistent volumes, which allow a user to attach the same volume to multiple pods running on different nodes. With the introduction of Dynamic Resource Allocation (DRA) we now have a way to allocate any type of resource with similar semantics.
In this talk, we discuss how DRA’s ability to allocate non-node-local resources has unlocked the potential to read / write remote GPU memory over high-bandwidth, multi-node NVLinks. We begin with an introduction on how DRA models non-node-local resources in general, followed by the specifics of how we have leveraged this capability to enable lightning fast multi-node training and inference on the NVIDIA GB200 NVL72 supercomputer. As part of this, we discuss how this support has been pushed to all major cloud providers and integrated with their managed Kubernetes offerings. We conclude with a demo.
World Congress 2025
July 10, 2025 · 13:30–14:00
Stage 5
Anshul Jindal, Martin Piercy
World Congress 2025
July 11, 2025 · 13:00–14:00
NVIDIA Lounge (Booth A16/17)
Anshul Jindal, Martin Piercy
World Congress 2025
July 10, 2025 · 10:00–11:00
NVIDIA Lounge (Booth A16/17)
Anshul Jindal, Kevin Klues, Martin Piercy
World Congress 2025
July 11, 2025 · 10:00–11:00
NVIDIA Lounge (Booth A16/17)
Miguel Martínez, Roman