World Congress 2024

Running AI Workloads in Containers and Kubernetes

July 18, 2024 16:10 โ€“ 16:40 ยท 30 min STAGE 8 (120)

What this session covers

Kubernetes is quickly becoming the platform of choice for running machine learning and AI workloads in the cloud. However, running these workloads efficiently poses unique challenges, from resource management to performance optimization.

In this talk, we dive into the details of how GPUs are made available to such workloads when running under Kubernetes. As part of this, we discuss various options for sharing GPUs between them. These techniques include simple time-slicing, MPS, and MIG.

By the end of this session, attendees will have a comprehensive understanding of how GPU support in Kubernetes works under the hood, as well as the knowledge required to make the most efficient use of GPUs in their own applications. We conclude with a demo.

Related talks at this congress

Open session

World Congress 2024

July 18, 2024 ยท 15:30โ€“16:00

STAGE 8 (120)

From foundation model to hosted AI solution in minutes

Kevin Klues, Stephan Schosser

Kevin Klues
Stephan Schosser
Open session

World Congress 2024

July 18, 2024 ยท 12:10โ€“12:40

STAGE 9 (600)

Containers and Kubernetes made easy: Deep dive into Podman Desktop and new AI capabilities

Stevan Le Meur

Stevan is Product Manager for Red Hat

Stevan Le Meur
Open session

World Congress 2024

July 18, 2024 ยท 10:50โ€“11:20

STAGE 9 (600)

Supercharge your cloud-native applications with Generative AI

Cedric Clyburn

Cedric Clyburn is Developer Advocate at Red Hat

Cedric Clyburn
Open session

World Congress 2024

July 19, 2024 ยท 11:00โ€“11:30

STAGE 4 (500)

Using Containers to deploy AI Models across our microscopy platform

Sebastian Rhode

Carl Zeiss Microscopy GmbH, Software Architect AI Solutions

Sebastian Rhode
All sessions at this congress