World Congress 2025
July 11, 2025 ยท 12:00โ14:00
M4 (40 Seats)
Model Compression Techniques for Efficient LLM Deployment
Harshita Seth, Lavinia Ghita, Sergio Perez, Ziv Ilan
World Congress 2025
As the adoption of LLMs continues to grow, the complexity of fine-tuning, and deploying these models has become a significant bottleneck. Manual processes and fragmented workflows can lead to errors, inconsistencies, and delays that hinder innovation and progress. This hands-on workshop introduces LLMOps, an approach to automating the entire LLM evaluation, and inference lifecycle using a GitOps-based methodology.
Participants will learn how to build an end-to-end automated pipeline leveraging NVIDIA NIMs and Nemo Microservices for fine-tuning, evaluation, and deployment of LLMs. Through practical demonstrations, we will explore how to ensure seamless integration, validation, and deployment of updates, leading to faster development cycles, improved accuracy, and increased reliability.
Key Topics: - Kubernetes-based LLM Pipelines - Argo CD for Continuous Delivery - Argo Workflows for LLM Workflow Automation - Cloud-Agnostic Deployment
World Congress 2025
July 11, 2025 ยท 12:00โ14:00
M4 (40 Seats)
Harshita Seth, Lavinia Ghita, Sergio Perez, Ziv Ilan
World Congress 2025
July 10, 2025 ยท 10:10โ10:40
Mainstage
Alejandro Saucedo
Director of Engineering, Science & Product at Zalando
World Congress 2025
July 11, 2025 ยท 12:15โ14:15
M6 (40 Seats)
Daniel Evenschor, Lars Langerbein, Matthias Kordt, Sebastian Klenke
World Congress 2025
July 10, 2025 ยท 10:30โ12:30
M3 (45 Seats)
Avin Zarlez
Staff SW Engineer & Developer Evangelist at Arm