> Markdown version of [/@ziv-ilan](https://www.wearedevelopers.com/@ziv-ilan). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Ziv Ilan GenAI Solution Architect at NVIDIA focusing on agentic workflows and LLM deployment. ## About Ziv Ilan works as a GenAI Solution Architect at NVIDIA, helping developers build and scale enterprise AI applications. He tackles the practical challenges of deploying large language models, focusing on agentic workflows and system reliability. His work guides teams in moving beyond basic AI prompts to construct specialized tools and end-to-end agents. He also breaks down model compression methods—like quantization, pruning, and knowledge distillation—to make LLM inference more efficient and reduce computational costs. ## Past Sessions ### World Congress 2025 · July 9, 2025 Berlin, Germany - [NVIDIA Expert Session: Best Practices and Techniques for Building Agentic and RAG applications](https://www.wearedevelopers.com/events/world-congress-2025/sessions/548-nvidia-expert) · 60 min - [NVIDIA Expert Session: Efficient Generative AI Inference Using Model Compression Techniques: Quantiz](https://www.wearedevelopers.com/events/world-congress-2025/sessions/575-nvidia-expert) · 60 min - [NVIDIA Expert Session: Sovereign AI in Practice: Building, Evaluating, and Scaling Multilingual LLMs](https://www.wearedevelopers.com/events/world-congress-2025/sessions/676-nvidia-expert) · 60 min - [Learn to Build Agentic AI Workflows for Enterprise Applications](https://www.wearedevelopers.com/events/world-congress-2025/sessions/730-learn-to-build) · 120 min - [ Model Compression Techniques for Efficient LLM Deployment](https://www.wearedevelopers.com/events/world-congress-2025/sessions/811-model-compression) · 120 min