> Markdown version of [/jobs/ext/2448236-sre-azure-cloud](https://www.wearedevelopers.com/jobs/ext/2448236-sre-azure-cloud). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SRE Azure Cloud - **Company:** Digitive LLC - **Location:** Pleasanton, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Query Performance, Application Performance Management, Microsoft Azure, Bash Shell, Cloud Computing, Databases, Continuous Integration, Data Recovery, DevOps, Distributed Systems, Monitoring of Systems, Python (Programming Language), Key Management, PostgreSQL, Log Analysis, Microsoft SQL Server, SQL Azure, MySQL, Performance Tuning, Windows PowerShell, Redis, Release Management, Reliability Engineering, Prometheus, Azure DevOps Pipelines, Software Deployment, Virtual Machines, Cloud Monitoring, Grafana, Mttr, Kubernetes Helm Charts, Software Troubleshooting, Git, Event Driven Architecture, Containerization, Kubernetes, Deployment Automation, Cosmos DB, Azure AKS, Bitbucket, Database Monitoring, Restful APIs, Dynatrace, Docker, Microservices - **Published:** August 28, 2026 - **Apply:** https://www.dice.com/job-detail/44c87538-aa26-46be-a06b-afde88b57713 ## About the Role Ideal Candidate: A hands-on Senior SRE with expertise in Stage/Production deployments, Azure-based microservices, observability (Grafana, Prometheus, Graylog, Azure Monitor), troubleshooting using Application Insights, database operations, and customer escalation management, focused on maintaining highly reliable and available production services., The ideal candidate will possess strong experience in cloud operations, microservices, production support, deployment automation, observability platforms, and database troubleshooting, with a proven ability to rapidly diagnose and resolve complex production issues., * Strong troubleshooting and analytical skills. * Production support and incident management expertise. * Customer-first mindset. * Excellent communication and stakeholder management. * Ability to perform effectively during critical outages and high-severity incidents. * Continuous improvement and automation mindset. ## Description We are seeking a highly skilled Senior Site Reliability Engineer (SRE) to ensure the reliability, availability, performance, and operational stability of customer-facing applications and services. This role is focused on Stage/Production deployments, Monitoring & Observability, Troubleshooting, Incident Management, and Customer Escalation support across Azure-based microservices environments., Production & Stage Operations * Manage and support Stage and Production environments. * Execute application, infrastructure, configuration, and database deployments. * Validate releases, perform health checks, and coordinate rollback activities. * Support change management and production readiness reviews. Monitoring & Observability * Build and maintain dashboards, alerts, and monitoring solutions. * Monitor application, infrastructure, and database health using logs, metrics, traces, and telemetry. * Improve observability coverage and reduce alert noise. * Proactively identify reliability and performance issues before customer impact. Troubleshooting & Incident Response * Troubleshoot software, infrastructure, configuration, deployment, and database-related issues. * Lead incident response activities and production recovery efforts. * Perform root cause analysis (RCA) and implement preventive actions. * Develop operational runbooks and troubleshooting documentation. Customer Escalation Management * Investigate and resolve customer-reported production issues. * Act as a technical lead during high-priority incidents. * Partner with Engineering, Product, and Customer Support teams to drive issue resolution. * Provide timely communication and status updates during major incidents. Required Technical Skills Cloud & Infrastructure: Microsoft Azure, Azure Kubernetes Service (AKS), Azure Virtual Machines, App Services, Azure Storage, Azure Networking, Application Gateway, Azure Key Vault Microservices & Containerization: Kubernetes, Docker, Helm Charts, Microservices Architecture, REST APIs, Event-Driven Architecture, Distributed Systems Troubleshooting CI/CD & DevOps: Azure DevOps Pipelines, Bitbucket, Git, Helm-based Deployments, CI/CD Release Management, Deployment Automation Monitoring & Observability: Grafana, Prometheus, Graylog, Azure Monitor, Application Insights, Log Analytics, Alerting & Dashboard Management, Distributed Tracing, SLI/SLO Monitoring Troubleshooting Expertise: Application Performance Issues, Production Incident Management, Configuration & Environment Issues, Deployment Failures & Rollbacks, Kubernetes & Container Troubleshooting, Network & Connectivity Issues, Root Cause Analysis (RCA) Databases: Azure SQL / SQL Server, PostgreSQL / MySQL, Cosmos DB, Redis, Query Performance Tuning, Database Monitoring, Backup & Recovery Automation & Scripting: PowerShell, Python, Bash Preferred Experience * Supporting enterprise SaaS applications in Production environments. * Azure-based microservices platforms running on AKS. * Customer-facing production support and escalation management. * 24x7 on-call and incident response environments. * Site Reliability Engineering (SRE) best practices including SLIs, SLOs, MTTR, and service availability management. ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [MySQL Protocol Features You Should Be Aware Of](https://www.wearedevelopers.com/videos/100267-mysql-protocol-features-you-should-be-aware-of) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Branch your database like your code: How schema changes and pull requests go hand in hand](https://www.wearedevelopers.com/videos/350-branch-your-database-like-your-code-how-schema-changes-and-pull-requests-go-hand-in-hand) - [Azure-Well Architected Framework - designing mission critical workloads in practice](https://www.wearedevelopers.com/videos/1529-azure-well-architected-framework-designing-mission-critical-workloads-in-practice) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)