> Markdown version of [/@lavinia-ghita](https://www.wearedevelopers.com/@lavinia-ghita). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lavinia Ghita Making massive AI models faster to deploy and building agentic workflows at NVIDIA. ## About Lavinia Ghita helps developers make generative AI run faster and cheaper. As a solution architect at NVIDIA, she focuses on taking massive language models and shrinking them down for practical enterprise deployment. Much of her work centers on the mechanics of model compression, including quantization, pruning, and knowledge distillation. She frequently guides hands-on engineers through building custom agentic AI workflows, evaluating system performance, and optimizing code for better GPU acceleration using CUDA. ## Past Sessions ### World Congress 2025 · July 9, 2025 Berlin, Germany - [NVIDIA Expert Session: Efficient Generative AI Inference Using Model Compression Techniques: Quantiz](https://www.wearedevelopers.com/events/world-congress-2025/sessions/575-nvidia-expert) · 60 min - [Learn to Build Agentic AI Workflows for Enterprise Applications](https://www.wearedevelopers.com/events/world-congress-2025/sessions/730-learn-to-build) · 120 min - [NVIDIA Expert Session: GPUs demystified - How to accelerate your code](https://www.wearedevelopers.com/events/world-congress-2025/sessions/793-nvidia-expert) · 60 min - [ Model Compression Techniques for Efficient LLM Deployment](https://www.wearedevelopers.com/events/world-congress-2025/sessions/811-model-compression) · 120 min - [NVIDIA Expert Session: CUDA Developer Best Practices](https://www.wearedevelopers.com/events/world-congress-2025/sessions/855-nvidia-expert) · 60 min - [NVIDIA Expert Session: Customize Generative AI Models for Industry-Specific Solutions](https://www.wearedevelopers.com/events/world-congress-2025/sessions/897-nvidia-expert) · 60 min