> Markdown version of [/events/world-congress-2025/sessions/811-model-compression](https://www.wearedevelopers.com/events/world-congress-2025/sessions/811-model-compression). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Model Compression Techniques for Efficient LLM Deployment - **Date:** Friday, Jul 11, 2025 - **Time:** 12:00–14:00 (120 min) - **Room:** M4 (40 Seats) - **Event:** World Congress 2025 - **Tags:** ai, deep learning, llms, machine learning, python ## Description The rapid advancements in large language models (LLMs) have unlocked transformative potential across industries, with state-of-the-art models excelling in popular AI reasoning benchmarks. However, their practical application often presents challenges due to their significant size, high computational costs, deployment complexities, and latency issues. To address these limitations, model compression techniques have emerged as effective solutions to reduce model size and computational requirements while preserving core capabilities. This tutorial will provide a comprehensive overview of key techniques, including quantization, pruning, and knowledge distillation. Participants will gain an understanding of the theoretical foundations behind each method, explore their use cases, and engage in practical, hands-on examples to solidify their learning. ## Speakers ### [Harshita Seth](https://www.wearedevelopers.com/@harshita-seth) Senior Solution Architect at Nvidia ### [Lavinia Ghita](https://www.wearedevelopers.com/@lavinia-ghita) NVIDIA, Solution Architect ### [Sergio Perez](https://www.wearedevelopers.com/@sergio-perez) Solution Architect at NVIDIA ### [Ziv Ilan](https://www.wearedevelopers.com/@ziv-ilan) Ziv Ilan, Solution Architect GenAI at NVIDIA ## Related talks at this congress - [Exploring LLMs across clouds](https://www.wearedevelopers.com/events/world-congress-2025/sessions/525-exploring-llms) — Tomislav Tipurić - [Adding knowledge to open-source LLMs](https://www.wearedevelopers.com/events/world-congress-2025/sessions/734-adding-knowledge-to) — Harshita Seth, Sergio Perez - [Inside the Mind of an LLM](https://www.wearedevelopers.com/events/world-congress-2025/sessions/831-inside-the-mind-of) — Emanuele Fabbiani - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/events/world-congress-2025/sessions/862-llmops-driven-fine) — Anshul Jindal