AI Observability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+7 more
Job description
Nebius is building a high-performance AI cloud platform, and we are looking for an AI Observability Engineer to own the AI observability backbone-from LLM and agent tracing to platform and infrastructure metrics-so the Azure AI platform is visible, measurable, reliable, and self-service., * Stand up and operate LLM and agent monitoring with Langfuse.
- Capture traces, latency, token usage, cost, quality scores, prompt and model-version analytics, and safety signals.
- Build lightweight internal tooling and exporters in Python.
- Design and maintain Grafana dashboards, Prometheus metrics, and the Azure observability stack.
- Instrument platform and AI workloads for health, usage, cost, and SLA reporting.
- Feed telemetry and operational insights into the Platform Engineering backlog.
- Own Terraform IaC and CI/CD for observability tooling.
- Support incident investigation and root-cause analysis.
Requirements
- 5-8 years of experience in observability, SRE, platform engineering, DevOps, or cloud engineering.
- Strong experience with Azure Monitor, Application Insights, Log Analytics, and Managed Grafana.
- Hands-on experience with Langfuse, Grafana, and Prometheus.
- Experience with Terraform and CI/CD.
- Python skills for instrumentation, exporters, and automation.
- Familiarity with ML workloads and AI-specific metrics.
- Knowledge of logs, metrics, traces, dashboards, alerting, SLIs, and SLOs.
- Intermediate or higher English.
Nice-to-haves:
- PromQL and Kusto Query Language.
- OpenTelemetry, including GenAI semantic conventions.
- LLM evaluation frameworks.
- AI cost dashboards and FinOps.
- Alerting, on-call, and incident management tooling.
- AKS and Kubernetes observability., Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire.
Benefits & conditions
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
What’s it like to work at Nebius:
Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI
About the company
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.adzuna.nlGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
MLOps And AI Driven Development
How to Become an AI Engineer
MLOps – What’s the deal behind it?
Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?