Observability & Evaluation Engineer

NTT DATA, Inc.
Charlotte, NC, United States
20 days ago
Apply on www.beyondcharlotte.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Automation of Tests Continuous Integration Python (Programming Language) Scaled Agile Framework Large Language Models Grafana Kubernetes Low Latency Machine Learning Operations Virtual Agents

Job description

Observability & Evaluation Engineer will build telemetry, tracing, dashboards, evaluation suites, alerts, service objectives, runbooks, and readiness evidence for Tachyon agent releases. This role ensures production AI systems can be monitored, evaluated, improved, and supported with clear operational visibility., Implement observability and telemetry for LLM-powered applications, agents, tools, and platform services.

Build evaluation suites for agent behavior, prompt quality, response quality, retrieval performance, latency, reliability, and safety signals.

Develop dashboards, alerts, traces, metrics, service objectives, and reporting for production readiness.

Work with platform engineers and Product Owners to define monitoring requirements and evaluation metrics.

Automate evidence collection for release readiness, operational reviews, and governance checkpoints.

Create runbooks and support documentation for priority agent releases.

Analyze production behavior and recommend improvements to reliability, performance, and quality.

Requirements

7+ years of engineering experience with observability, monitoring, test automation, platform operations, or AI/ML systems.

5+ years of strong hands-on Python experience.

5+ years of Experience with dashboards, metrics, alerts, traces, logs, SLOs, and production monitoring.

5+ years of Understanding of LLM evaluation, prompt evaluation, RAG evaluation, or AI quality assessment approaches.

5+ years of Experience working in Agile engineering teams and production support environments.

Required Skills / Knowledge

Python, telemetry, tracing, monitoring, dashboards, alerting, SLOs, evaluation frameworks, test automation, and production operations.

Understanding of LLMs, agents, RAG, prompt performance, retrieval quality, latency, and reliability metrics.

Experience with observability tools and open telemetry concepts.

Preferred Qualifications

Experience with GenAI observability, AI evaluation tools, ML monitoring, or platform reliability engineering.

Experience in regulated environments with evidence and readiness documentation.

Kubernetes, cloud platforms, and CI/CD experience.

Expected Outcomes

Operational dashboards and evaluation suites for priority agent releases.

Clear readiness evidence, alerts, SLOs, and runbooks.

Improved quality, reliability, and trust in production Agentic AI systems.

About the company

NTT DATA is a $30 billion trusted global innovator of business and technology services. We serve 75% of the Fortune Global 100 and are committed to helping clients innovate, optimize and transform for long term success. As a Global Top Employer, we have diverse experts in more than 50 countries and a robust partner ecosystem of established and start-up companies. Our services include business and technology consulting, data and artificial intelligence, industry solutions, as well as the development, implementation and management of applications, infrastructure and connectivity. We are one of the leading providers of digital and AI infrastructure in the world. NTT DATA is a part of NTT Group, which invests over $3.6 billion each year in R&D to help organizations and society move confidently and sustainably into the digital future. Visit us at us.nttdata.com

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.beyondcharlotte.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:05 min

Implementing monitoring and observability for AI software deployments

Alejandro Saucedo Alejandro Saucedo · World Congress 2025

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

5:48 min

Balancing delivery latency with stream reliability and scale

Phil Cluff · LIVE

2:12 min

Navigating technical clarity as a global black belt

Chris Heilmann +2 · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

Videos

See all

Related articles

See all