> Markdown version of [/jobs/ext/3126969-system-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3126969-system-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # System Reliability Engineer - **Company:** On-Demand Group - **Location:** United States - **Experience:** Expert - **Salary:** $187,200.0 - $228,800.0 - **Contract:** Permanent contract - **Skills:** Application Performance Management, Microsoft Azure, Cloud Computing, Cloud Engineering, Databases, Continuous Integration, Software Debugging, DevOps, Distributed Systems, Monitoring of Systems, Reliability Engineering, Subversion, Web Applications, Cloud Monitoring, Delivery Pipeline, Git, Kubernetes, Information Technology, Low Latency, Azure AKS, Software Version Control, Dynatrace, Devsecops - **Published:** September 28, 2026 - **Apply:** https://www.dice.com/job-detail/33ef5221-90ce-488d-ba9f-807753390b1a ## About the Role * Bachelor's degree in Computer Science, Management Information Sciences or area of functional responsibility preferred, or equivalent years of industry work experience * 5+ years in software or operations engineering * 2+ years of DevOps and Site Reliability engineering or similar experience with cloud-native solutions * Proven experience in DevOps culture and site reliability engineering focused on the customer, cross-functional autonomous teams, and continuous improvement * DevOps experience with a cloud-native web application hosted in Microsoft Azure * Familiarity with version control systems e.g., Git, SVN, CVS * Extensive database and operating systems experience * Experience in designing and implementing a continuous integration pipeline (CICD) * Experience in monitoring infrastructure, application uptime, latency, and performance on large distributed systems * Exhibit proficiency at troubleshooting various cloud and system related issues * Demonstrable cross-functional knowledge with systems, storage, networking, security, and databases * Excellent verbal and written communication skills to convey monitoring insights and collaborate across technical and non-technical teams. * Experience with cross-team collaboration. Partnering with DevOps/Platform Engineering, Production Support, and Architecture/Development teams to integrate monitoring solutions into existing applications, infrastructure, and automated pipelines Preferred Experience: * Kubernetes Monitoring: Proven experience with observability solutions tailored to Kubernetes environments (AKS), including monitoring containerized services and workload * Monitoring Solutions: Experience with designing, deploying, and maintaining monitoring frameworks using Dynatrace, Azure Monitor, and Application Insights, ensuring comprehensive visibility across distributed systems * Alerting & Incident Response: Configured alerting mechanisms based on proactive symptom-based thresholds, enabling rapid resolution of performance and reliability issues * Performance Analysis: Experience analyzing telemetry and monitoring data to identify bottlenecks in application and system performance within the Kubernetes ecosystem * A passion for leveraging observability tools to drive operational improvements in cloud-native applications ## Description The Sr. Site Reliability Engineer (SRE) is responsible for the availability, latency, performance, efficiency, monitoring, and emergency response. This role will be a member of a team that focuses on Support and SRE for the Digital Commerce Organization. The SRE drives continuous improvement in delivery of resilient, scalable, performant, secure, and high-quality services. Collaborating with DevOps, DevSecOps, and development teams, the SRE identifies cross-team issues which create risk for operations and resolving those issues with a mixture of engineering, troubleshooting expertise, and general operational guidance. Essential Functions: * Monitoring and Observability with Kubernetes-based applications and services running on Azure Kubernetes Services (AKS) * Plan, design, deploy, and operate Site Reliability Engineering capabilities for cloud products & services * Recognize and address sub-standard performance based on key performance indicators (KPIs) * Build monitoring that alerts on symptoms rather than outages * Strong debugging and problem-solving skills for complex, distributed systems. * Continuously build, automate, and improve upon capabilities that are secure, scalable, performant, and resilient * Work closely with Infrastructure, Network, Security, Architecture, and Development teams to build highly performing, scalable, and secure Azure environments * Define needs by documenting processes; includes research, planning and writing supporting documentation Additional Functions: In addition to the essential functions listed above, the incumbent may perform the following additional functions. * Participate in regulatory and compliance activities as necessary ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Azure-Well Architected Framework - designing mission critical workloads in practice](https://www.wearedevelopers.com/videos/1529-azure-well-architected-framework-designing-mission-critical-workloads-in-practice) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [DevSecOps: Injecting Security into Mobile CI/CD Pipelines](https://www.wearedevelopers.com/videos/273-devsecops-injecting-security-into-mobile-ci-cd-pipelines) - [The journey from developer to devops - what i've learnt along the way](https://www.wearedevelopers.com/videos/238-the-journey-from-developer-to-devops-what-i-ve-learnt-along-the-way) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering)