Azure Cloud DevOps Engineer

Tekshapers Inc
Minneapolis, MN, United States
3 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Starter
Experience required
1 year minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Continuous Integration Noise Reduction Linux Python (Programming Language) Pattern Recognition Runbook Software Vulnerability Management
+12 more
Google Cloud GitHub Copilot Office365 Mttr Cloudformation Kubernetes Terraform Splunk Dynatrace Devsecops Pagerduty Servicenow

Job description

  • Own Reliability Outcomes: Define and track SLIs/SLOs, manage error budgets, and drive continuous improvement to availability, latency, and resiliency for critical services.
  • Operate and Mature Observability/AIOps Platform: Build and tune monitoring, dashboards, alerting, and correlation (logs/metrics/traces) using tools like Dynatrace and Splunk to reduce noise and accelerate detection/diagnosis.
  • Lead Incident Response & Problem Management: Run on-call/war rooms, conduct root-cause analysis, publish post-incident reviews, and ensure corrective and preventive actions are delivered.
  • Automate Toil & Enable Self-Healing: Create runbooks, scripts, and workflows for automated remediation, safe changes, and guardrails to improve MTTR.
  • Partner with Engineering & Stakeholders: Consult on architecture, release readiness, capacity planning, and operational standards; translate reliability risks into executive-ready KPI updates (MTTD/MTTR, error budget burn, recurring toil).

Requirements

  • AIOps & SRE Fundamentals: 5+ years in SRE/production operations, including SLO/SLI, error budgets, incident management, and automated remediation patterns.
  • Observability Toolset: 3+ years building dashboards, alerts, and troubleshooting with tools such as Dynatrace and Splunk (log/metric/trace correlation, alert tuning, noise reduction).
  • Cloud & Platform Engineering: 3+ years operating services on AWS/Azure/Google Cloud Platform; strong expertise in Linux, networking, containers/Kubernetes, and IaC (e.g., Terraform/CloudFormation).
  • Automation & Agentic Ops Mindset: Proficient in Python/Bash and CI/CD; experienced in building runbooks, self-healing workflows, and participating in on-call rotations.
  • AI-Enabled SRE / Intelligent Ops: 1-2+ years applying AI-assisted incident response (auto-summarization, auto-triage, pattern detection) and predictive alerting/anomaly detection in production environments.
  • Security / DevSecOps: Working knowledge of vulnerability management, secrets/cert governance, and secure CI/CD gates.

Preferred Skills & Experience:

  • Release Safety & Resiliency Engineering: Experience with progressive delivery (canary/blue-green), chaos testing/DR drills, performance/capacity engineering, and well-architected reviews (e.g., Azure WARA/Azure Advisor).
  • ITSM & Tooling Integration: Hands-on experience integrating monitoring signals with ServiceNow and PagerDuty to build low-friction escalation workflows.

General Baseline Expectation:

  • Enterprise AI Tool Proficiency: Demonstrate consistent use (minimum 90% weekly usage) of enterprise-approved AI tools (e.g., GitHub Copilot, Microsoft 365 Copilot) to enhance coding, documentation, and overall delivery velocity.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:15 min

Introduction to artificial intelligence driven development

Natalie Pistunovich · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · World Congress 2021

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all