World Congress 2026 Europe
July 10, 2026 Β· 15:40β16:10
Stage 2
Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing
Piotr Zaniewski
Head of Engineering Enablement at vCluster Labs
World Congress 2026 Europe
Most of the LLM workloads now are deployed on Kubernetes clusters with GPU nodes and letβs be honest this is the most expensive resource in the cluster. Currently, using GPUs in passthrough mode locks a single model to an entire GPU, leading to severe underutilization (~30%). In this talk I will explain how to manage GPU resources in an efficient way and attendees will understand how GPU cards are configured in a Kubernetes cluster, what is the difference between the three main Nvidia GPU installation modes: Passthrough, vGPU and MIG and how everything works behind the scene. I will demonstrate how Multi instance GPUs are the best solution for packing LLMs on Kubernetes and how it should be used in an advanced case scenario like packing multiple LLMs in the same cluster sharing the same GPUs without causing the noisy neighbor problem.
World Congress 2026 Europe
July 10, 2026 Β· 15:40β16:10
Stage 2
Piotr Zaniewski
Head of Engineering Enablement at vCluster Labs
World Congress 2026 Europe
July 10, 2026 Β· 16:20β16:50
Stage 3 - powered by AWS
Jeremy Murray
Founder and CEO of Stack8s
World Congress 2026 Europe
July 10, 2026 Β· 14:45β16:45
Room M1 (60 Seats)
Duan Lightfoot
Senior Developer Advocate at Akamai
World Congress 2026 Europe
July 9, 2026 Β· 13:00β15:00
Room M2 (40 Seats)
Anshul Jindal, Mohak Chadha