TELECOMMUTE Sr. SRE Engineer (Azure Expert)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+13 more
Requirements
- Azure Administration and Architecture
- Networking and Connectivity
- Virtual Machines and Platform Services
- Azure Monitor and Log Analytics
- Azure Identity and Access Management
- Azure Networking, NSGs, Firewalls, Load Balancing
- Azure Backup and Disaster Recovery
Strong expertise in Azure networking (VNets, routing, firewalls, private links, load balancing).
Hands-on proficiency with infrastructure-as-code and automated deployments. (Must have Terraform and Git Enterprise, orchestration engines)
Exposure and understanding of building, deploying and managing API Gateways
Strong understanding of Azure security controls, governance, and compliance frameworks.
Full stack observability e.g. MELTS principles golden signals, and automation response using DataDog, New Relic, Splunk or other leading tools.
Strong FinOps expertise
Scripting skills (PowerShell, Bash, Python, React).
Strong understanding of Devops practices, tooling, and SDLC methods
Strong exposure to Anthropic, Open AI, platforms and associated tools & practices e.g Harness, Token usage, Skills, LLM and SLM concepts, Orchestration engines, and agent cost management.
Strong understanding of Site Reliability Engineering principles, including SLIs, SLOs, SLAs, error budgets, reliability targets, and service health measurement.
Experience designing observability strategies across metrics, logs, traces, synthetic monitoring, alerting, dashboards, and operational telemetry.
Ability to define actionable alerts that identify customer-impacting symptoms, reduce noise, and support rapid incident triage.
Proven capability in incident response, root cause analysis, blameless post-incident reviews, corrective action tracking, and operational learning.
Experience reducing toil through automation, self-service tooling, runbook automation, self-healing patterns, and repeatable engineering solutions.
Strong knowledge of capacity planning, performance engineering, load testing, scalability modelling, saturation analysis, and demand forecasting.
Experience with resilience validation techniques including chaos engineering, game days, failover testing, disaster recovery exercises, and operational readiness testing.
Ability to establish production readiness standards, reliability acceptance criteria, operational runbooks, service ownership models, and support handover practices.
Working knowledge of deployment reliability practices such as canary releases, blue-green deployments, rollback strategies, feature flags, and release health monitoring.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Mastering Remote Work: Tips for Developers
Is Software Engineering Over-Saturated?
Fully Remote Software Engineer Jobs