TELECOMMUTE Sr. SRE Engineer (Azure Expert)

Dminds Solutions
United States
3 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Microsoft Azure Bash Shell DevOps Disaster Recovery Python (Programming Language) Load Testing Log Analysis Routing Windows PowerShell Systems Development Life Cycle Reliability Engineering
+13 more
Virtual Machines Datadog Scripting Load Balancing ReactJS Large Language Models Firewalls (Computer Science) Git Deployment Automation Api Gateway Terraform Splunk New Relic (SaaS)

Requirements

  • Azure Administration and Architecture
  • Networking and Connectivity
  • Virtual Machines and Platform Services
  • Azure Monitor and Log Analytics
  • Azure Identity and Access Management
  • Azure Networking, NSGs, Firewalls, Load Balancing
  • Azure Backup and Disaster Recovery

Strong expertise in Azure networking (VNets, routing, firewalls, private links, load balancing).

Hands-on proficiency with infrastructure-as-code and automated deployments. (Must have Terraform and Git Enterprise, orchestration engines)

Exposure and understanding of building, deploying and managing API Gateways

Strong understanding of Azure security controls, governance, and compliance frameworks.

Full stack observability e.g. MELTS principles golden signals, and automation response using DataDog, New Relic, Splunk or other leading tools.

Strong FinOps expertise

Scripting skills (PowerShell, Bash, Python, React).

Strong understanding of Devops practices, tooling, and SDLC methods

Strong exposure to Anthropic, Open AI, platforms and associated tools & practices e.g Harness, Token usage, Skills, LLM and SLM concepts, Orchestration engines, and agent cost management.

Strong understanding of Site Reliability Engineering principles, including SLIs, SLOs, SLAs, error budgets, reliability targets, and service health measurement.

Experience designing observability strategies across metrics, logs, traces, synthetic monitoring, alerting, dashboards, and operational telemetry.

Ability to define actionable alerts that identify customer-impacting symptoms, reduce noise, and support rapid incident triage.

Proven capability in incident response, root cause analysis, blameless post-incident reviews, corrective action tracking, and operational learning.

Experience reducing toil through automation, self-service tooling, runbook automation, self-healing patterns, and repeatable engineering solutions.

Strong knowledge of capacity planning, performance engineering, load testing, scalability modelling, saturation analysis, and demand forecasting.

Experience with resilience validation techniques including chaos engineering, game days, failover testing, disaster recovery exercises, and operational readiness testing.

Ability to establish production readiness standards, reliability acceptance criteria, operational runbooks, service ownership models, and support handover practices.

Working knowledge of deployment reliability practices such as canary releases, blue-green deployments, rollback strategies, feature flags, and release health monitoring.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all