World Congress 2025 Aug 20, 2025 Session details

A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes

Kevin Klues

The NVIDIA GB200 joins 72 distributed GPUs into one supercomputer. Learn how the Compute Domain abstraction and Kubernetes DRA instantly solve the massive multi-node networking nightmare.

Pause
Mute Enter Fullscreen
#1 about 1 min

Supporting NVIDIA GB200 GPUs on Kubernetes

Adapting Kubernetes to support the unique architecture of the new GB200 GPU family.

#2 about 2 min

Scaling beyond single-node GPU communication limits

How multi-node NVLink uses NV switches to connect GPUs across separate physical machines.

#3 about 2 min

Kubernetes GPU operator limitations with GB200

Why out-of-the-box cluster configurations fail to provide secure shared memory access across multi-node setups.

#4 about 1 min

Enabling secure memory exchange with IMEX

How IMEX allows isolated GPU processes to directly read and write remote memory over NVLink.

#5 about 2 min

CUDA APIs for IMEX fabric attached memory

Mapping remote hardware through local pointers using specific CUDA memory handle exports.

#6 about 6 min

Four partitioning layers for IMEX access

Managing physical wiring, switch partitions, dynamic soft domains, and per-process channel permissions.

#7 about 5 min

Orchestrating dynamic IMEX domains for workloads

Configuring daemons to build secure point-to-point topologies mirroring Kubernetes pod deployments.

#8 about 5 min

Simplifying cluster setups using compute domains

Hiding multi-node communication wiring beneath custom API server objects and resource claim templates.

#9 about 3 min

Automating compute domains via the DRA driver

How controllers and kubelet plugins coordinate IMEX daemons and inject channel permissions sequentially.

#10 about 2 min

Routing cross-rack traffic seamlessly with NCCL

Binding segmented NVLink partitions automatically so higher-level libraries fallback to InfiniBand between isolated domains.

#11 about 7 min

Demonstrating multi-node high bandwidth memory transfers

Measuring memory read and write speeds by running an MPI workload across clustered compute resources.

#12 about 2 min

Cluster prerequisites to deploy compute domains

Enabling the required feature flags, CDIs, and custom NVIDIA daemons within your infrastructure.

Matching moments

1:37 min

Unlocking direct GPU access within managed Kubernetes platforms

Kevin Klues Kevin Klues

2:43 min

Bursting GPU capacity over hybrid Kubernetes networks

Jeremy Murray Jeremy Murray · WWC Europe 2026

2:59 min

Scaling up and scaling out GPU clusters

Michael Kagan Michael Kagan +1 · WWC Europe 2026

3:10 min

Scaling customized inference models with NVIDIA NIM

Anshul Jindal Anshul Jindal · WWC 2025

1:54 min

Scaling applications across multi-node GPU clusters and architectures

Paul Graham Paul Graham · LIVE

1:20 min

How the Kubernetes AI scheduler enables fractional GPU sharing

Piotr Zaniewski Piotr Zaniewski · WWC Europe 2026

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Trust, But Verify: Continuous GPU Validation at Scale

Kyle Bell

VP of AI @ TensorWave

Kyle Bell
Open session

World Congress 2026 North America

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

Compute for your AI model: GPUs, LPUs, TPUs and beyond..

Kushaagra Goyal

Tech Lead at Rubrik, ex-CTO at Gan.AI, ex-Databricks

Kushaagra Goyal
Open session

World Congress 2026 North America

Autonomous Infrastructure: Building AI Agents for Global-Scale Capacity Efficiency

Tommy Tran, Gregoire Colin

Tommy Tran
Gregoire Colin
Open session

World Congress 2026 North America

Run your agents in Kubernetes: Build once, deploy anywhere. But really?

Michal Salanci

Senior Systems Engineer at ESET Cybersecurity

Michal Salanci
Open session

World Congress 2026 North America

From Cloud Native to Multi-Cloud Native: Write Once, Deploy Anywhere

Sandeep Pal

Principal Member of Technical Staff at Salesforce

Sandeep Pal