World Congress 2026 Europe

Compress, Cut, and Distill: The Latest Gen AI Model Compression Techniques in Practice

July 10, 2026 12:15 – 14:15 · 120 min Room M2 (40 Seats)

What this session covers

This training lab explores the art and engineering of compressing large language models to make them cheaper, faster, and easier to deploy while preserving practical capability. Designed for a broad audience that spans beginners to advanced practitioners, the workshop will introduce foundational concepts for newcomers, share implementation patterns and pitfalls for experienced engineers, and highlight cutting-edge research directions for specialists.

Workshop Preparation: - Please bring your own laptop. - Please review the following document and prepare accordingly before the workshop: https://developer.nvidia.com/dli/getready

Related talks at this congress

Open session

World Congress 2026 Europe

July 9, 2026 · 13:25–13:30

Airstream 1

Smaller Voice Models

Sohaib Ahmad

CEO of Neuphonic

Sohaib Ahmad
Open session

World Congress 2026 Europe

July 10, 2026 · 16:20–16:50

Stage 6 - powered by Microsoft

Fine-Tuning Small Language Models for Agentic AI

Björn Buchhold

Technology Evangelist at CID

Björn Buchhold
Open session

World Congress 2026 Europe

July 9, 2026 · 14:10–14:40

Stage 6 - powered by Microsoft

Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated

Christin Pohl

Global Black Belt Solution Engineer at Microsoft

Christin Pohl
Open session

World Congress 2026 Europe

July 9, 2026 · 10:30–12:30

Room M2 (40 Seats)

Teaching AI to Code in Every Language with NVIDIA NeMo

Antonio Rueda-Toicen, Marco Gullotto DE, Nikita Pavlichenko

Antonio Rueda-Toicen
Marco Gullotto DE
Nikita Pavlichenko
All sessions at this congress