System Reliability Engineer

On-Demand Group
United States
3 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Compensation
$187,200.0 - $228,800.0
Working hours
Regular working hours
Job source

Tech stack

Application Performance Management Microsoft Azure Cloud Computing Cloud Engineering Databases Continuous Integration Software Debugging DevOps Distributed Systems Monitoring of Systems Reliability Engineering Subversion
+11 more
Web Applications Cloud Monitoring Delivery Pipeline Git Kubernetes Information Technology Low Latency Azure AKS Software Version Control Dynatrace Devsecops

Job description

The Sr. Site Reliability Engineer (SRE) is responsible for the availability, latency, performance, efficiency, monitoring, and emergency response. This role will be a member of a team that focuses on Support and SRE for the Digital Commerce Organization. The SRE drives continuous improvement in delivery of resilient, scalable, performant, secure, and high-quality services. Collaborating with DevOps, DevSecOps, and development teams, the SRE identifies cross-team issues which create risk for operations and resolving those issues with a mixture of engineering, troubleshooting expertise, and general operational guidance.

Essential Functions:

  • Monitoring and Observability with Kubernetes-based applications and services running on Azure Kubernetes Services (AKS)

  • Plan, design, deploy, and operate Site Reliability Engineering capabilities for cloud products & services
  • Recognize and address sub-standard performance based on key performance indicators (KPIs)
  • Build monitoring that alerts on symptoms rather than outages

  • Strong debugging and problem-solving skills for complex, distributed systems.
  • Continuously build, automate, and improve upon capabilities that are secure, scalable, performant, and resilient
  • Work closely with Infrastructure, Network, Security, Architecture, and Development teams to build highly performing, scalable, and secure Azure environments
  • Define needs by documenting processes; includes research, planning and writing supporting documentation

Additional Functions: In addition to the essential functions listed above, the incumbent may perform the following additional functions.

  • Participate in regulatory and compliance activities as necessary

Requirements

  • Bachelor’s degree in Computer Science, Management Information Sciences or area of functional responsibility preferred, or equivalent years of industry work experience
  • 5+ years in software or operations engineering
  • 2+ years of DevOps and Site Reliability engineering or similar experience with cloud-native solutions
  • Proven experience in DevOps culture and site reliability engineering focused on the customer, cross-functional autonomous teams, and continuous improvement
  • DevOps experience with a cloud-native web application hosted in Microsoft Azure
  • Familiarity with version control systems e.g., Git, SVN, CVS
  • Extensive database and operating systems experience
  • Experience in designing and implementing a continuous integration pipeline (CICD)
  • Experience in monitoring infrastructure, application uptime, latency, and performance on large distributed systems
  • Exhibit proficiency at troubleshooting various cloud and system related issues
  • Demonstrable cross-functional knowledge with systems, storage, networking, security, and databases
  • Excellent verbal and written communication skills to convey monitoring insights and collaborate across technical and non-technical teams.

  • Experience with cross-team collaboration. Partnering with DevOps/Platform Engineering, Production Support, and Architecture/Development teams to integrate monitoring solutions into existing applications, infrastructure, and automated pipelines

Preferred Experience:

  • Kubernetes Monitoring: Proven experience with observability solutions tailored to Kubernetes environments (AKS), including monitoring containerized services and workload

  • Monitoring Solutions: Experience with designing, deploying, and maintaining monitoring frameworks using Dynatrace, Azure Monitor, and Application Insights, ensuring comprehensive visibility across distributed systems

  • Alerting & Incident Response: Configured alerting mechanisms based on proactive symptom-based thresholds, enabling rapid resolution of performance and reliability issues

  • Performance Analysis: Experience analyzing telemetry and monitoring data to identify bottlenecks in application and system performance within the Kubernetes ecosystem

  • A passion for leveraging observability tools to drive operational improvements in cloud-native applications

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:15 min

Key lessons learned from implementing automated mobile DevSecOps

Moataz Nabil Moataz Nabil · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all