SRE / AIOps Engineer

ITBMS Inc.
Minneapolis, MN, United States
3 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Bash Shell Continuous Integration Noise Reduction Linux Disaster Recovery Python (Programming Language) Pattern Recognition Performance Tuning Reliability Engineering
+20 more
Cloud Services Runbook Software Vulnerability Management Scripting Google Cloud Performance Testing GitHub Copilot Office365 Grafana Mttr Cloudformation AI Platforms Kubernetes Low Latency Terraform Splunk Dynatrace Devsecops Pagerduty Servicenow

Job description

  • Define and track SLIs/SLOs, manage error budgets, and drive continuous improvements in availability, latency, and resiliency.
  • Build and optimize observability and AIOps platforms, including monitoring, dashboards, alerting, and log/metric/trace correlation.
  • Work with tools such as Dynatrace and Splunk to reduce alert noise and accelerate incident detection and troubleshooting.
  • Lead incident response, on-call activities, war rooms, and root-cause analysis.
  • Conduct post-incident reviews and ensure corrective and preventive actions are completed.
  • Develop runbooks, scripts, and automated remediation workflows to reduce operational toil and improve MTTR.
  • Partner with Engineering and business stakeholders on architecture, release readiness, capacity planning, and operational standards.
  • Translate reliability metrics and risks into executive-level reporting, including MTTD, MTTR, error-budget burn, and recurring toil.

Requirements

  • 5+ years of experience in SRE / Production Operations.
  • Strong understanding of SLOs, SLIs, error budgets, incident management, and automated remediation.
  • 3+ years of hands-on experience with observability tools such as Dynatrace and Splunk.
  • Strong experience with logs, metrics, traces, alert tuning, dashboard development, and noise reduction.
  • 3+ years operating cloud-based services on AWS, Azure, or Google Cloud Platform.
  • Strong knowledge of Linux, networking, containers, Kubernetes, and Infrastructure as Code.
  • Experience with Terraform, CloudFormation, or similar IaC technologies.
  • Strong scripting/automation skills using Python and/or Bash.
  • Experience with CI/CD pipelines and automated operational workflows.
  • Experience building runbooks, self-healing workflows, and remediation automation.
  • Hands-on experience with AI-assisted incident response, including auto-triage, incident summarization, pattern detection, anomaly detection, or predictive alerting.
  • Working knowledge of Security/DevSecOps, vulnerability management, secrets/certificate governance, and secure CI/CD practices.

Preferred Skills

  • Experience with canary and blue-green deployments.
  • Experience with chaos engineering, disaster recovery drills, performance testing, and capacity planning.
  • Experience with Well-Architected Reviews, Azure WARA, or Azure Advisor.
  • Hands-on integration with ServiceNow and PagerDuty.
  • Experience building low-friction incident escalation and notification workflows.
  • Experience with AI tools such as GitHub Copilot, Microsoft 365 Copilot, or other enterprise-approved AI platforms.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · World Congress 2021

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all