World Congress 2025
July 11, 2025 · 12:00–14:00
M4 (40 Seats)
Model Compression Techniques for Efficient LLM Deployment
Harshita Seth, Lavinia Ghita, Sergio Perez, Ziv Ilan
World Congress 2025
While Large Language Models are remarkably powerful, their knowledge is typically limited to the data available during training. This makes their applications challenging when business contexts, or cultural nuances change or are not well represented in their original training data. In many cases, in-context learning techniques such as retrieval-augmented generation (RAG) can help to bridge this gap by providing models with relevant information at runtime. However, when a model lacks foundational understanding of domain-specific knowledge, use case, or local cultural context, even advanced retrieval methods may fail. In this session, NVIDIA experts will explain how to enrich language models with new knowledge, expanding their capabilities in specialized business, engineering, or scientific domains, and adjusting adaptation to new languages, cultures, and values.
World Congress 2025
July 11, 2025 · 12:00–14:00
M4 (40 Seats)
Harshita Seth, Lavinia Ghita, Sergio Perez, Ziv Ilan
World Congress 2025
July 11, 2025 · 12:00–14:00
M7 (18 Seats)
Miguel Martínez, Roman
World Congress 2025
July 10, 2025 · 14:50–15:20
Stage 6 - Red Hat
Cedric Clyburn, Roberto Carratalá
World Congress 2025
July 11, 2025 · 13:00–13:30
Stage 7
Emanuele Fabbiani
Head of AI at xtream