World Congress 2026 Europe Jul 10, 2026 Session details

Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing

Piotr Zaniewski

Stop choosing between GPU utilization and Kubernetes stability. Discover how combining vCluster and NVIDIA KAI creates instant, multi-tenant AI sandboxes. Run different schedulers safely and maximize hardware density.

Pause
Mute Enter Fullscreen
#1 about 4 min

Challenges of shared Kubernetes clusters for AI workloads

Shared global components complicate software upgrades and limit scheduling flexibility for AI workloads.

#2 about 2 min

How the Kubernetes AI scheduler enables fractional GPU sharing

The dynamic resource allocator divides graphics hardware between specific workloads to improve utilization.

#3 about 2 min

Isolating control planes with virtual Kubernetes clusters

Running the application programming interface server inside a pod reduces the operational blast radius.

#4 about 3 min

Navigating the spectrum of Kubernetes multi-tenancy models

Virtual clusters balance full infrastructure isolation against sharing global services like ingress controllers.

#5 about 1 min

Exploring the internal architecture of vcluster components

The sinker component synchronizes specific resources between virtual environments and the underlying host platform.

#6 about 2 min

Verifying hardware access and exploring AI inference scaling

Large-scale machine learning processing requires robust cluster management software to efficiently allocate computational power.

#7 about 2 min

Demonstrating GPU workloads by generating text haikus

Generating haikus via a deployed language model verifies that the cluster correctly handles GPU-accelerated workloads.

#8 about 2 min

Configuring virtual clusters to run custom AI schedulers

Defining custom configurations encapsulates specialized scheduling logic within an isolated tenant environment.

#9 about 3 min

Deploying and connecting to the sandboxed control plane

Using terminal applications verifies the correct installation and functionality of the isolated control plane pods.

#10 about 2 min

Allocating fractional GPU resources using the isolated scheduler

Applying custom configuration files allows individual pods to consume exact percentages of shared graphics hardware.

#11 about 2 min

Provisioning multiple virtual clusters for concurrent tenant configurations

Deploying simultaneous isolated environments satisfies conflicting software version requirements across different engineering teams.

#12 about 2 min

Scaling virtual clusters to maximize hardware cost savings

Running thousands of encapsulated sandboxes on shared host nodes drastically reduces organizational computing costs.

#13 about 3 min

Verifying concurrent versions of isolated AI schedulers

Checking cluster stateful sets confirms that multiple scheduler versions operate independently without software interference.

#14 about 1 min

Documentation and community resources for cluster management tools

Directing developers to official documentation and community channels provides support for troubleshooting complex integrations.

#15 about 3 min

Configuring virtual clusters for multi-node GPU parallelization

Targeting dedicated infrastructure allows specialized workloads to distribute execution across multiple discrete computing devices.

Matching moments

1:54 min

Speaker background and open source Kubernetes edge computing projects

Gaurav Gahlot Gaurav Gahlot · World Congress 2026 Europe

4:49 min

Structuring compute and data services for AI models

Radu Vunvulea Radu Vunvulea · World Congress 2025

2:04 min

Optimizing and deploying containerized AI inference workloads

Ankit Patel Ankit Patel · World Congress 2024

1:37 min

Unlocking direct GPU access within managed Kubernetes platforms

Kevin Klues Kevin Klues

5:07 min

Actions Runner Controller architecture and cluster isolation

Bassem Dghaidi · LIVE

1:30 min

Overlooked AI infrastructure and operational deployment barriers

Stan Girard Stan Girard · World Congress 2024

Upcoming sessions on this topic

Open session

World Congress 2026 North America

September 25, 2026 · 15:30–16:00

Stage 7

Trust, But Verify: Continuous GPU Validation at Scale

Kyle Bell

VP of AI @ TensorWave

Kyle Bell
Open session

World Congress 2026 North America

September 24, 2026 · 11:00–11:30

Stage 1

Application-Defined Compute: Rethinking Infrastructure for AI Applications

Anurag Goel

Founder and CEO of Render

Anurag Goel
Open session

World Congress 2026 North America

September 25, 2026 · 09:00–09:30

Stage 2

Autonomous Infrastructure: Building AI Agents for Global-Scale Capacity Efficiency

Tommy Tran

Software Engineer at Meta

Tommy Tran
Open session

World Congress 2026 North America

September 25, 2026 · 11:00–11:30

Stage 5

Managing GPUs by Just Asking, Infrastructure in the Age of MCP

Jessica Garson Beauchemin

Developer Relations Lead, Community at Runpod

Jessica Garson Beauchemin
Open session

World Congress 2026 North America

September 23, 2026 · 10:45–12:45

Stage 10

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Sr. AI Engineer, Akamai

Duan Lightfoot
Open session

World Congress 2026 North America

September 25, 2026 · 13:30–14:00

Mainstage

A Hands-On Developer Guide to Inference Engineering

Ankit Patel, Philip Kiely

Ankit Patel
Philip Kiely