> Markdown version of [/jobs/ext/3137542-sre-aiops-engineer](https://www.wearedevelopers.com/jobs/ext/3137542-sre-aiops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SRE / AIOps Engineer - **Company:** ITBMS Inc. - **Location:** Minneapolis, MN, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Bash Shell, Continuous Integration, Noise Reduction, Linux, Disaster Recovery, Python (Programming Language), Pattern Recognition, Performance Tuning, Reliability Engineering, Cloud Services, Runbook, Software Vulnerability Management, Scripting, Google Cloud, Performance Testing, GitHub Copilot, Office365, Grafana, Mttr, Cloudformation, AI Platforms, Kubernetes, Low Latency, Terraform, Splunk, Dynatrace, Devsecops, Pagerduty, Servicenow - **Published:** September 29, 2026 - **Apply:** https://www.dice.com/job-detail/e0e9e5ce-0b8d-459e-ab01-19c552e5a80e ## About the Role * 5+ years of experience in SRE / Production Operations. * Strong understanding of SLOs, SLIs, error budgets, incident management, and automated remediation. * 3+ years of hands-on experience with observability tools such as Dynatrace and Splunk. * Strong experience with logs, metrics, traces, alert tuning, dashboard development, and noise reduction. * 3+ years operating cloud-based services on AWS, Azure, or Google Cloud Platform. * Strong knowledge of Linux, networking, containers, Kubernetes, and Infrastructure as Code. * Experience with Terraform, CloudFormation, or similar IaC technologies. * Strong scripting/automation skills using Python and/or Bash. * Experience with CI/CD pipelines and automated operational workflows. * Experience building runbooks, self-healing workflows, and remediation automation. * Hands-on experience with AI-assisted incident response, including auto-triage, incident summarization, pattern detection, anomaly detection, or predictive alerting. * Working knowledge of Security/DevSecOps, vulnerability management, secrets/certificate governance, and secure CI/CD practices. Preferred Skills * Experience with canary and blue-green deployments. * Experience with chaos engineering, disaster recovery drills, performance testing, and capacity planning. * Experience with Well-Architected Reviews, Azure WARA, or Azure Advisor. * Hands-on integration with ServiceNow and PagerDuty. * Experience building low-friction incident escalation and notification workflows. * Experience with AI tools such as GitHub Copilot, Microsoft 365 Copilot, or other enterprise-approved AI platforms. ## Description * Define and track SLIs/SLOs, manage error budgets, and drive continuous improvements in availability, latency, and resiliency. * Build and optimize observability and AIOps platforms, including monitoring, dashboards, alerting, and log/metric/trace correlation. * Work with tools such as Dynatrace and Splunk to reduce alert noise and accelerate incident detection and troubleshooting. * Lead incident response, on-call activities, war rooms, and root-cause analysis. * Conduct post-incident reviews and ensure corrective and preventive actions are completed. * Develop runbooks, scripts, and automated remediation workflows to reduce operational toil and improve MTTR. * Partner with Engineering and business stakeholders on architecture, release readiness, capacity planning, and operational standards. * Translate reliability metrics and risks into executive-level reporting, including MTTD, MTTR, error-budget burn, and recurring toil. ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Reducing Cognitive Overload Through Platform Engineering](https://www.wearedevelopers.com/videos/679-reducing-cognitive-overload-through-platform-engineering) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)