> Markdown version of [/@sergio-perez](https://www.wearedevelopers.com/@sergio-perez). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sergio Perez Optimizes LLM training and inference as a Senior Solution Architect at NVIDIA. ## About Sergio Perez optimizes how large language models run in production. As a Senior Solution Architect at NVIDIA, he focuses on conversational AI, specializing in quantization and inference serving to make massive models deployable. His work tackles the practical realities of generative AI for developers. He regularly teaches model compression, custom fine-tuning workflows, and how to build efficient retrieval-augmented generation systems for enterprise environments. Before NVIDIA, he engineered AI systems at Graphcore and Amazon. He holds a Ph.D. in computational fluid dynamics from Imperial College London. ## Past Sessions ### World Congress 2026 Europe 路 July 8, 2026 Berlin, Germany - [Physical AI for the Next Wave of Industrial Digitalisation](https://www.wearedevelopers.com/videos/100039-physical-ai-for-the-next-wave-of-industrial-digitalisation) 路 30 min 路 馃帴 Watch recording - [Nemotron: NVIDIA's open model strategy for developers](https://www.wearedevelopers.com/videos/100064-nemotron-nvidia-s-open-model-strategy-for-developers) 路 30 min 路 馃帴 Watch recording - [Compress, Cut, and Distill: The Latest Gen AI Model Compression Techniques in Practice](https://www.wearedevelopers.com/events/world-congress-2026-europe/sessions/1286-compress-cut-and) 路 120 min ### World Congress 2025 路 July 9, 2025 Berlin, Germany - [NVIDIA Expert Session: Efficient Generative AI Inference Using Model Compression Techniques: Quantiz](https://www.wearedevelopers.com/events/world-congress-2025/sessions/575-nvidia-expert) 路 60 min - [NVIDIA Expert Session: Sovereign AI in Practice: Building, Evaluating, and Scaling Multilingual LLMs](https://www.wearedevelopers.com/events/world-congress-2025/sessions/676-nvidia-expert) 路 60 min - [Adding knowledge to open-source LLMs](https://www.wearedevelopers.com/videos/1522-adding-knowledge-to-open-source-llms) 路 30 min 路 馃帴 Watch recording - [ Model Compression Techniques for Efficient LLM Deployment](https://www.wearedevelopers.com/events/world-congress-2025/sessions/811-model-compression) 路 120 min - [NVIDIA Expert Session: Customize Generative AI Models for Industry-Specific Solutions](https://www.wearedevelopers.com/events/world-congress-2025/sessions/897-nvidia-expert) 路 60 min ## Videos - [Nemotron: NVIDIA's open model strategy for developers](https://www.wearedevelopers.com/videos/100064-nemotron-nvidia-s-open-model-strategy-for-developers) 路 Sergio Perez - [Physical AI for the Next Wave of Industrial Digitalisation](https://www.wearedevelopers.com/videos/100039-physical-ai-for-the-next-wave-of-industrial-digitalisation) 路 Sergio Perez - [Adding knowledge to open-source LLMs](https://www.wearedevelopers.com/videos/1522-adding-knowledge-to-open-source-llms) 路 Harshita Seth, Sergio Perez ## Links - [LinkedIn](https://www.linkedin.com/in/sergiopp/)