> Markdown version of [/jobs/ext/2737077-observability-evaluation-engineer-nc](https://www.wearedevelopers.com/jobs/ext/2737077-observability-evaluation-engineer-nc). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Observability & Evaluation Engineer (NC) - **Company:** Diverse Lynx LLC - **Location:** Charlotte, NC, United States - **Experience:** Expert - **Salary:** $145,600.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Automation of Tests, Microsoft Azure, Continuous Integration, Python (Programming Language), Prometheus, Runbook, Datadog, Data Logging, Large Language Models, Grafana, Generative AI, AI Platforms, Kubernetes, Low Latency, Virtual Agents, Splunk, Dynatrace, Api Management, Docker, Microservices - **Published:** September 5, 2026 - **Apply:** https://www.careerjet.com/job/us30c4c44c0180db5b73c9c0d5ce47b13b/eaa ## About the Role * Strong Python development experience. * Hands-on experience with LLM, Generative AI, or AI agent evaluation. * Experience developing automated test and evaluation frameworks. * Strong knowledge of telemetry, metrics, logging, and distributed tracing. * Experience with monitoring dashboards and alerting. * Understanding of SLOs, SLIs, and production operations. * Experience analyzing prompt and model performance. * Strong troubleshooting and production support experience. Preferred Skills * OpenTelemetry * Prometheus / Grafana * Datadog / Splunk * LLM observability and evaluation platforms * AI agent and multi-step workflow monitoring * CI/CD * AWS, Azure, or GCP * Kubernetes / Docker * API testing and microservices * SRE and incident management practices, ABM Industries is seeking an Electrical Field Test Technician (NETA III / IV or equivalent experience) to join our Electrical Power Services team. The Electrical Field Test Technic… + 2 months ago ## Description We are seeking an experienced Observability & Evaluation Engineer to build and operationalize observability, telemetry, and evaluation capabilities for LLM applications and AI agent systems. The engineer will implement end-to-end tracing, metrics, dashboards, automated evaluation suites, SLOs, runbooks, and production-readiness evidence for priority AI and agent releases. The ideal candidate will have strong experience with LLM/agent evaluation, observability, Python, test automation, production operations, and model/prompt performance analysis., * Implement telemetry, logging, metrics, and distributed tracing for LLM applications and AI agents. * Build dashboards to monitor AI application health, model behavior, agent performance, latency, errors, and usage. * Develop automated LLM and AI agent evaluation suites for quality, reliability, safety, and performance. * Create reusable Python-based test automation and evaluation frameworks. * Analyze prompt and model performance and identify optimization opportunities. * Define and monitor SLIs, SLOs, operational metrics, and alerting for AI services. * Develop runbooks for troubleshooting, incident response, and production support. * Establish production-readiness criteria and maintain release-readiness evidence for priority agent releases. * Integrate evaluation and observability checks into CI/CD pipelines. * Troubleshoot production issues involving model performance, latency, availability, failures, and unexpected agent behavior. * Collaborate with AI/ML engineers, SRE, platform, product, and governance teams. ## Related Videos - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Beyond the Benchmark: How to Evaluate AI Agents in the Real World](https://www.wearedevelopers.com/videos/100269-beyond-the-benchmark-how-to-evaluate-ai-agents-in-the-real-world) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [With AIs wide open - WeAreDevelopers at All Things Open 2025](https://www.wearedevelopers.com/magazine/641-with-ais-wide-open-wearedevelopers-at-all-things-open-2025)