Observability & Evaluation Engineer (NC)

Diverse Lynx LLC
Charlotte, NC, United States
21 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$145,600.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Automation of Tests Microsoft Azure Continuous Integration Python (Programming Language) Prometheus Runbook Datadog Data Logging Large Language Models Grafana
+10 more
Generative AI AI Platforms Kubernetes Low Latency Virtual Agents Splunk Dynatrace Api Management Docker Microservices

Job description

We are seeking an experienced Observability & Evaluation Engineer to build and operationalize observability, telemetry, and evaluation capabilities for LLM applications and AI agent systems. The engineer will implement end-to-end tracing, metrics, dashboards, automated evaluation suites, SLOs, runbooks, and production-readiness evidence for priority AI and agent releases. The ideal candidate will have strong experience with LLM/agent evaluation, observability, Python, test automation, production operations, and model/prompt performance analysis., * Implement telemetry, logging, metrics, and distributed tracing for LLM applications and AI agents.

  • Build dashboards to monitor AI application health, model behavior, agent performance, latency, errors, and usage.
  • Develop automated LLM and AI agent evaluation suites for quality, reliability, safety, and performance.
  • Create reusable Python-based test automation and evaluation frameworks.
  • Analyze prompt and model performance and identify optimization opportunities.
  • Define and monitor SLIs, SLOs, operational metrics, and alerting for AI services.
  • Develop runbooks for troubleshooting, incident response, and production support.
  • Establish production-readiness criteria and maintain release-readiness evidence for priority agent releases.
  • Integrate evaluation and observability checks into CI/CD pipelines.
  • Troubleshoot production issues involving model performance, latency, availability, failures, and unexpected agent behavior.
  • Collaborate with AI/ML engineers, SRE, platform, product, and governance teams.

Requirements

  • Strong Python development experience.
  • Hands-on experience with LLM, Generative AI, or AI agent evaluation.
  • Experience developing automated test and evaluation frameworks.
  • Strong knowledge of telemetry, metrics, logging, and distributed tracing.
  • Experience with monitoring dashboards and alerting.
  • Understanding of SLOs, SLIs, and production operations.
  • Experience analyzing prompt and model performance.
  • Strong troubleshooting and production support experience.

Preferred Skills

  • OpenTelemetry
  • Prometheus / Grafana
  • Datadog / Splunk
  • LLM observability and evaluation platforms
  • AI agent and multi-step workflow monitoring
  • CI/CD
  • AWS, Azure, or GCP
  • Kubernetes / Docker
  • API testing and microservices
  • SRE and incident management practices, ABM Industries is seeking an Electrical Field Test Technician (NETA III / IV or equivalent experience) to join our Electrical Power Services team. The Electrical Field Test Technic…
  • 2 months ago

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:50 min

Introduction and the value of runbooks

Hila Fish · World Congress 2023

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

1:32 min

Structuring automated incident workflows between runbooks and raw models

Aram Hakobyan Aram Hakobyan +1 · World Congress 2026 Europe

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all