Site Reliability Engineer

Intersources Inc.
United States
10 days ago
Apply on www2.jobdiva.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$25,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Continuous Integration Reliability Engineering Prometheus Mesos Large Language Models Grafana Kubernetes Virtual Agents BIG-IP Access Policy Manager (APM) ArcSight Event Correlation Splunk
+2 more
Pagerduty Servicenow

Job description

Design observability stacks tailored for AI agent performance (latency, cost, quality). Implement anomaly detection for runtime errors, hallucinations, and agent drifts. Collaborate with Ops/SRE Agent to automate remediation workflows. Define reliability SLIs/SLOs for agent-driven systems. Architect and operationalize end-to-end observability frameworks (metrics, traces, logs, golden signals) across clusters, workloads, and services. Shape the orchestration platform roadmap for resiliency, scalability, and operational intelligence in alignment with business objectives.

Requirements

Background in SRE for AI systems or large distributed platforms. Strong with OpenTelemetry, Prometheus, APM Tools, Grafana, Splunk. Familiarity with AI observability (LLM trace monitoring, token cost tracking, drift detection). Ability to integrate AI reliability checks into CI/CD and production environments. Deep expertise in orchestration platforms (Kubernetes, Nomad, Mesos, or equivalent) at enterprise scale. Preferred Experience in AIOps or ML observability. Background in incident management (PagerDuty, OpsGenie, ServiceNow). Proven success architecting and delivering AIOps and NoOps solutions - including event correlation, AI-driven automation, and self-healing operations. Experience automating/programming in Python, Go, or similar, with experience building ML- or AI-integrated pipelines.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www2.jobdiva.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:06 min

Empowering site reliability engineers with integrated AI agents

Osmar Matos Osmar Matos · World Congress 2026 Europe

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:29 min

Evaluating orchestration tools for distributed container deployments

Aleksandr Kalikov · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

Videos

See all

Related articles

See all