World Congress 2026 Europe

GPU is not Monolithic : Packing LLMs with MIGs on Kubernetes

July 9, 2026 10:50 – 11:20 Β· 30 min Stage 8 - powered by Red Hat

What this session covers

Most of the LLM workloads now are deployed on Kubernetes clusters with GPU nodes and let’s be honest this is the most expensive resource in the cluster. Currently, using GPUs in passthrough mode locks a single model to an entire GPU, leading to severe underutilization (~30%). In this talk I will explain how to manage GPU resources in an efficient way and attendees will understand how GPU cards are configured in a Kubernetes cluster, what is the difference between the three main Nvidia GPU installation modes: Passthrough, vGPU and MIG and how everything works behind the scene. I will demonstrate how Multi instance GPUs are the best solution for packing LLMs on Kubernetes and how it should be used in an advanced case scenario like packing multiple LLMs in the same cluster sharing the same GPUs without causing the noisy neighbor problem.

Related talks at this congress

Open session

World Congress 2026 Europe

July 10, 2026 Β· 15:40–16:10

Stage 2

Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing

Piotr Zaniewski

Head of Engineering Enablement at vCluster Labs

Piotr Zaniewski
Open session

World Congress 2026 Europe

July 10, 2026 Β· 16:20–16:50

Stage 3 - powered by AWS

Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬

Jeremy Murray

Founder and CEO of Stack8s

Jeremy Murray
Open session

World Congress 2026 Europe

July 10, 2026 Β· 14:45–16:45

Room M1 (60 Seats)

Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs

Duan Lightfoot

Senior Developer Advocate at Akamai

Duan Lightfoot
Open session

World Congress 2026 Europe

July 9, 2026 Β· 13:00–15:00

Room M2 (40 Seats)

Accelerating AI Inference at Scale: A Deep Dive Into NVIDIA Dynamo on Kubernetes

Anshul Jindal, Mohak Chadha

Anshul Jindal
Mohak Chadha
All sessions at this congress