> Markdown version of [/jobs/ext/2944988-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2944988-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Intersources Inc. - **Location:** United States - **Salary:** $25,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Continuous Integration, Reliability Engineering, Prometheus, Mesos, Large Language Models, Grafana, Kubernetes, Virtual Agents, BIG-IP Access Policy Manager (APM), ArcSight Event Correlation, Splunk, Pagerduty, Servicenow - **Published:** September 16, 2026 - **Apply:** https://www2.jobdiva.com/portal/?a=62jdnw10t7d77ytz0lbaey9qxonk3b05a9tqw96rb7z6hlfo7xj79l9g6mp6aj2o&compid=0/jobs/31127648#/jobs/31127648 ## About the Role Background in SRE for AI systems or large distributed platforms. Strong with OpenTelemetry, Prometheus, APM Tools, Grafana, Splunk. Familiarity with AI observability (LLM trace monitoring, token cost tracking, drift detection). Ability to integrate AI reliability checks into CI/CD and production environments. Deep expertise in orchestration platforms (Kubernetes, Nomad, Mesos, or equivalent) at enterprise scale. Preferred Experience in AIOps or ML observability. Background in incident management (PagerDuty, OpsGenie, ServiceNow). Proven success architecting and delivering AIOps and NoOps solutions - including event correlation, AI-driven automation, and self-healing operations. Experience automating/programming in Python, Go, or similar, with experience building ML- or AI-integrated pipelines. ## Description Design observability stacks tailored for AI agent performance (latency, cost, quality). Implement anomaly detection for runtime errors, hallucinations, and agent drifts. Collaborate with Ops/SRE Agent to automate remediation workflows. Define reliability SLIs/SLOs for agent-driven systems. Architect and operationalize end-to-end observability frameworks (metrics, traces, logs, golden signals) across clusters, workloads, and services. Shape the orchestration platform roadmap for resiliency, scalability, and operational intelligence in alignment with business objectives. ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Keycloak case study: Making users happy with service level indicators and observability](https://www.wearedevelopers.com/videos/1599-keycloak-case-study-making-users-happy-with-service-level-indicators-and-observability) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)