Sergio Perez
Optimizes LLM training and inference as a Senior Solution Architect at NVIDIA.
Sergio Perez optimizes how large language models run in production. As a Senior Solution Architect at NVIDIA, he focuses on conversational AI, specializing in quantization and inference serving to make massive models deployable.
His work tackles the practical realities of generative AI for developers. He regularly teaches model compression, custom fine-tuning workflows, and how to build efficient retrieval-augmented generation systems for enterprise environments.
Before NVIDIA, he engineered AI systems at Graphcore and Amazon. He holds a Ph.D. in computational fluid dynamics from Imperial College London.