World Congress 2025

Model Compression Techniques for Efficient LLM Deployment

July 11, 2025 12:00 – 14:00 · 120 min M4 (40 Seats)
ai llms machine learning python deep learning

What this session covers

The rapid advancements in large language models (LLMs) have unlocked transformative potential across industries, with state-of-the-art models excelling in popular AI reasoning benchmarks. However, their practical application often presents challenges due to their significant size, high computational costs, deployment complexities, and latency issues. To address these limitations, model compression techniques have emerged as effective solutions to reduce model size and computational requirements while preserving core capabilities.

This tutorial will provide a comprehensive overview of key techniques, including quantization, pruning, and knowledge distillation. Participants will gain an understanding of the theoretical foundations behind each method, explore their use cases, and engage in practical, hands-on examples to solidify their learning.

Related talks at this congress

Open session

World Congress 2025

July 10, 2025 · 10:50–11:20

Stage 11

Exploring LLMs across clouds

Tomislav Tipurić

Chief Technology Officer, Nephos

Tomislav Tipurić
Open session

World Congress 2025

July 11, 2025 · 09:40–10:10

Stage 2

Adding knowledge to open-source LLMs

Harshita Seth, Sergio Perez

Harshita Seth
Sergio Perez
Open session

World Congress 2025

July 11, 2025 · 13:00–13:30

Stage 7

Inside the Mind of an LLM

Emanuele Fabbiani

Head of AI at xtream

Emanuele Fabbiani
Open session

World Congress 2025

July 11, 2025 · 14:20–14:50

Stage 5

LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices

Anshul Jindal

Sr. Solution Architect at NVIDIA

Anshul Jindal
All sessions at this congress