> Markdown version of [/events/world-congress-2026-europe/sessions/1012-gpu-is-not](https://www.wearedevelopers.com/events/world-congress-2026-europe/sessions/1012-gpu-is-not). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # GPU is not Monolithic : Packing LLMs with MIGs on Kubernetes - **Date:** Thursday, Jul 9, 2026 - **Time:** 10:50–11:20 (30 min) - **Room:** Stage 8 - powered by Red Hat - **Event:** World Congress 2026 Europe ## Description Most of the LLM workloads now are deployed on Kubernetes clusters with GPU nodes and let's be honest this is the most expensive resource in the cluster. Currently, using GPUs in passthrough mode locks a single model to an entire GPU, leading to severe underutilization (~30%). In this talk I will explain how to manage GPU resources in an efficient way and attendees will understand how GPU cards are configured in a Kubernetes cluster, what is the difference between the three main Nvidia GPU installation modes: Passthrough, vGPU and MIG and how everything works behind the scene. I will demonstrate how Multi instance GPUs are the best solution for packing LLMs on Kubernetes and how it should be used in an advanced case scenario like packing multiple LLMs in the same cluster sharing the same GPUs without causing the noisy neighbor problem. ## Speaker ### [Hajed Khlifi](https://www.wearedevelopers.com/@hajed-khlifi) AI / HPC Architect at Quantori ## Related talks at this congress - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/events/world-congress-2026-europe/sessions/1373-instant-kai) — Piotr Zaniewski - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/events/world-congress-2026-europe/sessions/1387-running-secure-life) — Jeremy Murray - [Agents That Own Their Inference: Building Production AI Agents on Dedicated GPUs](https://www.wearedevelopers.com/events/world-congress-2026-europe/sessions/1350-agents-that-own) — Duan Lightfoot - [Accelerating AI Inference at Scale: A Deep Dive Into NVIDIA Dynamo on Kubernetes](https://www.wearedevelopers.com/events/world-congress-2026-europe/sessions/1061-accelerating-ai) — Anshul Jindal, Mohak Chadha