Lavinia Ghita

Lavinia Ghita

Making massive AI models faster to deploy and building agentic workflows at NVIDIA.

Lavinia Ghita helps developers make generative AI run faster and cheaper. As a solution architect at NVIDIA, she focuses on taking massive language models and shrinking them down for practical enterprise deployment.

Much of her work centers on the mechanics of model compression, including quantization, pruning, and knowledge distillation. She frequently guides hands-on engineers through building custom agentic AI workflows, evaluating system performance, and optimizing code for better GPU acceleration using CUDA.