> Markdown version of [/videos/2126-from-black-box-to-glass-box-bedrock-agentcore-observability](https://www.wearedevelopers.com/videos/2126-from-black-box-to-glass-box-bedrock-agentcore-observability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # From Black Box to Glass Box : Bedrock AgentCore Observability Why did your AI hallucinate? Discover how Bedrock Agent Core uses automated tracing to expose every reasoning step, turning complex LLM logic into a transparent glass box. - **Speakers:** [Yasemin Aktürk](https://www.wearedevelopers.com/@yasemin-akturk) - **Event:** World Congress 2026 Europe - Virtual Stage - **Published:** July 7, 2026 - **Duration:** 21:15 - **URL:** https://www.wearedevelopers.com/videos/2126-from-black-box-to-glass-box-bedrock-agentcore-observability ## Summary Amazon Bedrock Agent Core provides the production-ready building blocks—like runtime, memory, observability, and evaluation—necessary to turn complex AI logic from a "black box into a glass box." In this demonstration, a customer support agent built with the AWS open-source Python SDK illustrates how developers can track exactly what happens during a user request. The architecture utilizes short-term memory for session state alongside an asynchronous model that extracts long-term semantic preferences, giving the agent persistent context without requiring custom code for memory management. By leveraging AWS CloudWatch observability, developers can inspect every reasoning step, tool call, and memory operation across different execution environments (including local tools, gateways, and external servers). Automated traces reveal granular performance metrics, such as individual API call latencies and LLM token usage, validating why internal tools execute faster than gateway-backed endpoints. The most powerful diagnostic capability showcased is the platform's automated evaluation metric, which scores responses based on correctness, goal success rate, and tool selection accuracy. For instance, tracing an incorrect agent response back to a missing warranty-check API call yielded a low 0.5 correctness score, instantly highlighting the factual error and LLM hallucination. This transparent logging ensures that every model decision is visible, enabling engineers to confidently refine prompts and improve agent logic without guessing. **Keywords:** amazon bedrock agent core, AWS cloudwatch observability, AI agent memory extraction, short-term session state, long-term semantic memory, aws open-source agent SDK, agent core gateway, tool selection accuracy, LLM hallucination debugging, AI reasoning traces, built-in correctness evaluation, API latency metrics, LLM token usage tracking, automated evaluation metrics, genAI observability dashboard ## Chapters 1. **Introduction to Amazon Bedrock Agent Core capabilities** (00:02) — Agent Core provides production-ready building blocks to simplify the overall development and infrastructure of AI agents. 1. **How Agent Core runtime executes and tracks agent requests** (01:02) — Runtime infrastructure paths prompts through sandbox environments to safely execute tool calls and track latency. 1. **Retaining conversation context with short and long-term memory** (03:47) — Background models automatically extract and synchronize important facts across both active and historical conversation memory streams. 1. **Architecture overview of the demo customer support agent** (05:57) — The agent requires access to specific endpoints and environment tools to successfully fulfill customer support queries. 1. **Navigating the GenAI observability dashboard in Amazon CloudWatch** (07:28) — Centralized monitoring dashboards present detailed traces, event memory metrics, and token usage statistics for root cause analysis. 1. **Analyzing tool selection and execution within trace details** (10:41) — Trace trajectories demonstrate whether an agent routed user requests correctly while simultaneously tracking total token usage. 1. **Reviewing traces to uncover missed tool calls and inaccuracies** (15:28) — Trace logs expose factual inaccuracies when agents hallucinate answers instead of querying appropriate external data tools. 1. **Evaluating agent response accuracy and tool selection logic** (17:22) — Built-in correctness metrics score factual accuracy to help developers diagnose the root cause of hallucinated answers. 1. **Reviewing agent interactions completed without external tool execution** (20:00) — Execution steps verify that basic conversational queries successfully bypass external endpoints to lower overall tool usage. ## Related Moments - [Capturing AI session telemetry for deep usage insights](https://www.wearedevelopers.com/videos/100238-how-building-with-ai-can-double-the-throughput-of-your-engineering-team) (from "How building with AI can double the throughput of your engineering team") - [Building observability to verify unexpected AI agent behavior](https://www.wearedevelopers.com/videos/100291-building-apis-for-agents-vs-systems-is-mcp-the-answer) (from "Building APIs for Agents vs Systems. Is MCP the answer?") - [Executing and analyzing the core AI agent sub-workflow](https://www.wearedevelopers.com/videos/1523-the-ai-agent-path-to-prod-building-for-reliability) (from "The AI Agent Path to Prod: Building for Reliability") - [Architectural building blocks for enterprise agent development platforms](https://www.wearedevelopers.com/videos/1538-composable-intelligence-how-henkel-and-microsoft-are-shaping-the-agent-ecosystem) (from "Composable Intelligence: How Henkel and Microsoft Are Shaping the Agent Ecosystem") - [Tracing agentic capabilities and step-by-step code execution](https://www.wearedevelopers.com/videos/100086-unlocking-the-ai-black-box-building-trust-in-the-era-of-agentic-production) (from "Unlocking the AI Black Box: Building Trust in the Era of Agentic Production") - [Evaluating unguided agent tools against context-aware investigation pipelines](https://www.wearedevelopers.com/videos/100308-beyond-chat-ai-workflows-that-actually-investigate-alerts-so-you-don-t-have-to-know-everything) (from "Beyond Chat: AI Workflows That Actually Investigate Alerts (So You Don't Have To Know Everything)") ## Related Articles - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Introducing Redis Agent Memory Server](https://www.wearedevelopers.com/magazine/699-introducing-redis-agent-memory-server) - [Never delegate the understanding](https://www.wearedevelopers.com/magazine/749-never-delegate-the-understanding) - [Liuba Gonta and Yuliya Khadasevic - GitHub Copilot Beyond the Basics - 10 Ways to Elevate Your Coding](https://www.wearedevelopers.com/magazine/490-liuba-gonta-and-yuliya-khadasevic-github-copilot-beyond-the-basics-10-ways-to-elevate-your-coding) ## Related Jobs - [Principal Product Manager, Agent Platform](https://www.wearedevelopers.com/jobs/ext/277541-principal-product-manager-agent-platform) at **GitHub** - [Senior AI Agent Software Engineer (Go, Python) (m/f/x)](https://www.wearedevelopers.com/jobs/48277-senior-ai-agent-software-engineer-go-python-m-f-x) at **Dynatrace** - [Staff Software Engineer, Copilot Experiences](https://www.wearedevelopers.com/jobs/ext/164361-staff-software-engineer-copilot-experiences) at **GitHub** - [AI Software Engineer (Germany)](https://www.wearedevelopers.com/jobs/48317-ai-software-engineer-germany) at **Sunhat** - [Senior Backend Developer — AI: MCP & Agent Engine](https://www.wearedevelopers.com/jobs/48297-senior-backend-developer-ai-mcp-agent-engine) at **basebox GmbH** - [Principal Field Architect - AI Agents](https://www.wearedevelopers.com/jobs/ext/1442858-principal-field-architect-ai-agents) at **Twilio**