Observability & Evaluation Engineer (NC)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+10 more
Job description
We are seeking an experienced Observability & Evaluation Engineer to build and operationalize observability, telemetry, and evaluation capabilities for LLM applications and AI agent systems. The engineer will implement end-to-end tracing, metrics, dashboards, automated evaluation suites, SLOs, runbooks, and production-readiness evidence for priority AI and agent releases. The ideal candidate will have strong experience with LLM/agent evaluation, observability, Python, test automation, production operations, and model/prompt performance analysis., * Implement telemetry, logging, metrics, and distributed tracing for LLM applications and AI agents.
- Build dashboards to monitor AI application health, model behavior, agent performance, latency, errors, and usage.
- Develop automated LLM and AI agent evaluation suites for quality, reliability, safety, and performance.
- Create reusable Python-based test automation and evaluation frameworks.
- Analyze prompt and model performance and identify optimization opportunities.
- Define and monitor SLIs, SLOs, operational metrics, and alerting for AI services.
- Develop runbooks for troubleshooting, incident response, and production support.
- Establish production-readiness criteria and maintain release-readiness evidence for priority agent releases.
- Integrate evaluation and observability checks into CI/CD pipelines.
- Troubleshoot production issues involving model performance, latency, availability, failures, and unexpected agent behavior.
- Collaborate with AI/ML engineers, SRE, platform, product, and governance teams.
Requirements
- Strong Python development experience.
- Hands-on experience with LLM, Generative AI, or AI agent evaluation.
- Experience developing automated test and evaluation frameworks.
- Strong knowledge of telemetry, metrics, logging, and distributed tracing.
- Experience with monitoring dashboards and alerting.
- Understanding of SLOs, SLIs, and production operations.
- Experience analyzing prompt and model performance.
- Strong troubleshooting and production support experience.
Preferred Skills
- OpenTelemetry
- Prometheus / Grafana
- Datadog / Splunk
- LLM observability and evaluation platforms
- AI agent and multi-step workflow monitoring
- CI/CD
- AWS, Azure, or GCP
- Kubernetes / Docker
- API testing and microservices
- SRE and incident management practices, ABM Industries is seeking an Electrical Field Test Technician (NETA III / IV or equivalent experience) to join our Electrical Power Services team. The Electrical Field Test Technic…
- 2 months ago
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
Dev Digest 121 - AI goes offline
Dev Digest 120 - Apple and peers
Dev Digest 132 - Binging WADFlix?