> Markdown version of [/videos/1862-building-agents-securely-at-scale-alfonso-graziano?t=1507](https://www.wearedevelopers.com/videos/1862-building-agents-securely-at-scale-alfonso-graziano?t=1507). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Building Agents Securely at Scale - Alfonso Graziano Alfonso Graziano warns that deploying AI agents blindly is a security nightmare. Stop relying on simplistic tutorials. Learn to implement hard guardrails and continuous evaluations for production-ready LLMs. - **Speakers:** [Alfonso Graziano](https://www.wearedevelopers.com/@alfonso-graziano) - **Event:** Coffee With Developers - **Published:** April 27, 2026 - **Duration:** 31:49 - **URL:** https://www.wearedevelopers.com/videos/1862-building-agents-securely-at-scale-alfonso-graziano ## Summary Moving from simplistic tutorials to production-ready AI agents exposes significant security and reliability gaps. Alfonso Graziano details how handling natural language queries and the non-deterministic nature of large language models requires a fundamental shift toward "AI native engineering." In this new paradigm, developers evolve from writing granular code to functioning as tech leads who manage, review, and orchestrate a team of automated agent workflows. A core challenge is safely deploying these agents within enterprise environments involving restricted permissions and sensitive knowledge bases. Graziano emphasizes that security cannot be treated as an afterthought or solved entirely by the LLM. He points to the OWASP Top 10 for agentic applications as a vital framework to defend against emerging vulnerabilities like prompt injection, while cautioning against unrestricted execution without hard guardrails. Maintaining a strict human-in-the-loop safety net is presented as a mandatory requirement before allowing an agent to perform high-impact or potentially destructive actions. To harden these systems for actual customers, engineering teams must collaborate explicitly with subject matter experts to construct comprehensive "golden datasets" that capture highly varied edge cases. Implementing automated, continuous evaluations (evals) and tracking failure modes using advanced observability mechanisms allows teams to systematically reduce hallucinations. Overall, bridging the gap between prototype and scalable integration relies on analyzing real user interactions and expert trace annotations, establishing a closed feedback loop that iteratively toughens the underlying models and agents. **Keywords:** ai agents, ai native engineering, golden datasets, owasp top 10, prompt injection, security guardrails, human in the loop, continuous evaluations, agentic design patterns, ai system observability, large language models, prompt engineering, non-deterministic outputs, trace annotations, ai software deployment ## Chapters 1. **Building client-facing AI agents for engineering teams** (00:25) — Practical implementations of AI agents focus primarily on serving internal engineering teams and end customers. 1. **Moving beyond simple prompts to reliable agentic systems** (01:50) — Constructing robust AI agents requires recognizing that natural language interfaces must handle unpredictable user queries effectively. 1. **Why simplistic AI agent tutorials fail in production** (04:05) — Most introductory guides ignore critical components like automated evaluations, golden datasets, and continuous user feedback loops. 1. **Implementing security guardrails and OWASP principles for agents** (05:47) — Applying appropriate access controls and addressing new vulnerabilities like indirect prompt injection protects non-deterministic applications. 1. **Essential resources for understanding agentic design and evaluation** (08:00) — Foundational knowledge of large language models and structured frameworks aids in building easily testable AI systems. 1. **Evolving developer roles into tech leads for AI agents** (09:55) — Software engineers must review generated outputs and confidently guide parallel agent workflows instead of blindly trusting automated code. 1. **Risks of granting AI agents complete system permissions** (13:27) — Running experimental agents on personal machines exposes sensitive credentials and local files to unexpected behaviors. 1. **Mitigating hallucinations and sycophancy in tool-calling agents** (16:32) — Restricting tool access and applying framework-level guardrails helps prevent deployed agents from inventing fictitious functions. 1. **Building golden datasets and feedback loops for reliability** (18:37) — Gathering real user interactions and expert annotations forms the foundation for continuously evaluating and improving agent performance. 1. **Refining system prompts to eliminate specific failure modes** (23:40) — Adjusting domain-specific instructions within the prompt configuration significantly boosts evaluation scores and inherently prevents common errors. 1. **Practical enterprise use cases for automating complex workflows** (25:07) — Deploying intelligent agents for complex data search and reproducing repetitive development tasks safely accelerates team productivity. 1. **Learning resources and community engagement for AI engineers** (27:13) — Specialized courses, upcoming literature, and developer conferences offer structured approaches for mastering advanced software integration capabilities. ## Related Moments - [Preserving engineering fundamentals in agentic development](https://www.wearedevelopers.com/videos/1897-agents-version-control-and-bunnies-daniel-siegl-david-payr) (from "Agents, Version Control and Bunnies - Daniel Siegl & David Payr") - [Rethinking team structures around AI agent capabilities](https://www.wearedevelopers.com/videos/1539-agentic-devops-how-ai-powered-automation-transforms-software-delivery-on-github-and-azure) (from "Agentic DevOps: How AI-Powered Automation Transforms Software Delivery on GitHub and Azure") - [Security integration and AI skepticism in developer tooling](https://www.wearedevelopers.com/videos/1830-wearedevelopers-live-speculaitions) (from "WeAreDevelopers LIVE - SpeculAItions") - [Identifying emerging security vulnerabilities in generative AI agents](https://www.wearedevelopers.com/videos/1383-the-state-of-genai-machine-learning-in-2025) (from "The State of GenAI & Machine Learning in 2025") - [Fusing developer experience and platform engineering for agentic SDLC](https://www.wearedevelopers.com/videos/100266-ai-won-t-fix-your-engineering-culture) (from "AI Won't Fix Your Engineering Culture") - [Designing agentic AI solutions for the enterprise](https://www.wearedevelopers.com/videos/1831-ai-for-enterprise-developers-dr-damir-dobric) (from "AI for Enterprise Developers - Dr. Damir Dobric") ## Related Articles - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Exploring AI: Opportunities and Risks for Developers](https://www.wearedevelopers.com/magazine/522-exploring-ai-opportunities-and-risks-for-developers) - [WWC24 Talk - Scott Hanselman - AI: Superhero or Supervillain?](https://www.wearedevelopers.com/magazine/469-wwc24-talk-scott-hanselman-ai-superhero-or-supervillain) ## Related Jobs - [Senior AI Agent Software Engineer (Go, Python) (m/f/x)](https://www.wearedevelopers.com/jobs/48277-senior-ai-agent-software-engineer-go-python-m-f-x) at **Dynatrace** - [AI Software Engineer (Germany)](https://www.wearedevelopers.com/jobs/48317-ai-software-engineer-germany) at **Sunhat** - [Principal Product Manager, Agent Platform](https://www.wearedevelopers.com/jobs/ext/277541-principal-product-manager-agent-platform) at **GitHub** - [AI Operations Manager (all genders)](https://www.wearedevelopers.com/jobs/48263-ai-operations-manager-all-genders) at **envelio** - [AI Full Stack Engineer](https://www.wearedevelopers.com/jobs/ext/1354435-ai-full-stack-engineer) at **Almedia** - [Staff Software Engineer, Copilot Experiences](https://www.wearedevelopers.com/jobs/ext/164361-staff-software-engineer-copilot-experiences) at **GitHub**