> Markdown version of [/jobs/ext/2598247-observability-evaluation-engineer](https://www.wearedevelopers.com/jobs/ext/2598247-observability-evaluation-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Observability & Evaluation Engineer - **Company:** NTT DATA, Inc. - **Location:** Charlotte, NC, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Automation of Tests, Continuous Integration, Python (Programming Language), Scaled Agile Framework, Large Language Models, Grafana, Kubernetes, Low Latency, Machine Learning Operations, Virtual Agents - **Published:** August 16, 2026 - **Apply:** https://www.beyondcharlotte.com/job.asp?id=3355986215&tx=KJ3939FFF&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role 7+ years of engineering experience with observability, monitoring, test automation, platform operations, or AI/ML systems. 5+ years of strong hands-on Python experience. 5+ years of Experience with dashboards, metrics, alerts, traces, logs, SLOs, and production monitoring. 5+ years of Understanding of LLM evaluation, prompt evaluation, RAG evaluation, or AI quality assessment approaches. 5+ years of Experience working in Agile engineering teams and production support environments. Required Skills / Knowledge Python, telemetry, tracing, monitoring, dashboards, alerting, SLOs, evaluation frameworks, test automation, and production operations. Understanding of LLMs, agents, RAG, prompt performance, retrieval quality, latency, and reliability metrics. Experience with observability tools and open telemetry concepts. Preferred Qualifications Experience with GenAI observability, AI evaluation tools, ML monitoring, or platform reliability engineering. Experience in regulated environments with evidence and readiness documentation. Kubernetes, cloud platforms, and CI/CD experience. Expected Outcomes Operational dashboards and evaluation suites for priority agent releases. Clear readiness evidence, alerts, SLOs, and runbooks. Improved quality, reliability, and trust in production Agentic AI systems. ## Description Observability & Evaluation Engineer will build telemetry, tracing, dashboards, evaluation suites, alerts, service objectives, runbooks, and readiness evidence for Tachyon agent releases. This role ensures production AI systems can be monitored, evaluated, improved, and supported with clear operational visibility., Implement observability and telemetry for LLM-powered applications, agents, tools, and platform services. Build evaluation suites for agent behavior, prompt quality, response quality, retrieval performance, latency, reliability, and safety signals. Develop dashboards, alerts, traces, metrics, service objectives, and reporting for production readiness. Work with platform engineers and Product Owners to define monitoring requirements and evaluation metrics. Automate evidence collection for release readiness, operational reviews, and governance checkpoints. Create runbooks and support documentation for priority agent releases. Analyze production behavior and recommend improvements to reliability, performance, and quality. ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Swapping Low Latency Data Storage Under High Load](https://www.wearedevelopers.com/videos/746-swapping-low-latency-data-storage-under-high-load) - [Beyond the Benchmark: How to Evaluate AI Agents in the Real World](https://www.wearedevelopers.com/videos/100269-beyond-the-benchmark-how-to-evaluate-ai-agents-in-the-real-world) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Unlocking the AI Black Box: Building Trust in the Era of Agentic Production](https://www.wearedevelopers.com/videos/100086-unlocking-the-ai-black-box-building-trust-in-the-era-of-agentic-production) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)