World Congress 2026 Europe

Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing

July 10, 2026 15:40 – 16:10 Β· 30 min Stage 2

What this session covers

Kubernetes offers many ways to share GPUs, but a single, cluster-wide scheduler often forces trade-offs between utilization, stability, and team autonomy. This talk shows how vCluster makes the NVIDIA Kubernetes AI Scheduler (KAI) run as an opt-in service for each tenantβ€”so platform teams can raise GPU density while keeping operations predictable.

What We’ll Cover

Problem statement – why mixed workloads leave GPUs under-used and complicate on-call.
vCluster fundamentals – lightweight control planes that isolate scheduling logic, not hardware.
KAI at a glance – fractional GPU allocation, gang queues, topology awareness.
Live demonstration – two vClusters on one host

Key Takeaways

A reproducible pattern for running different schedulers side-by-side.
Practical steps to increase GPU utilisation without adding more clusters.
An isolation model that lets teams experiment safely.

Related talks at this congress

Open session

World Congress 2026 Europe

July 10, 2026 Β· 16:20–16:50

Stage 3 - powered by AWS

Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬

Jeremy Murray

Founder and CEO of Stack8s

Jeremy Murray
Open session

World Congress 2026 Europe

July 9, 2026 Β· 10:50–11:20

Stage 8 - powered by Red Hat

GPU is not Monolithic : Packing LLMs with MIGs on Kubernetes

Hajed Khlifi

AI / HPC Architect at Quantori

Hajed Khlifi
Open session

World Congress 2026 Europe

July 9, 2026 Β· 15:30–16:00

Stage 5

Building a Cloud Platform Where Everything is Just Another Kubernetes Resource

Patrick Koss

Tech Lead at STACKIT

Patrick Koss
Open session

World Congress 2026 Europe

July 9, 2026 Β· 13:00–15:00

Room M2 (40 Seats)

Accelerating AI Inference at Scale: A Deep Dive Into NVIDIA Dynamo on Kubernetes

Anshul Jindal, Mohak Chadha

Anshul Jindal
Mohak Chadha
All sessions at this congress