System Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+11 more
Job description
The Sr. Site Reliability Engineer (SRE) is responsible for the availability, latency, performance, efficiency, monitoring, and emergency response. This role will be a member of a team that focuses on Support and SRE for the Digital Commerce Organization. The SRE drives continuous improvement in delivery of resilient, scalable, performant, secure, and high-quality services. Collaborating with DevOps, DevSecOps, and development teams, the SRE identifies cross-team issues which create risk for operations and resolving those issues with a mixture of engineering, troubleshooting expertise, and general operational guidance.
Essential Functions:
-
Monitoring and Observability with Kubernetes-based applications and services running on Azure Kubernetes Services (AKS)
- Plan, design, deploy, and operate Site Reliability Engineering capabilities for cloud products & services
- Recognize and address sub-standard performance based on key performance indicators (KPIs)
-
Build monitoring that alerts on symptoms rather than outages
- Strong debugging and problem-solving skills for complex, distributed systems.
- Continuously build, automate, and improve upon capabilities that are secure, scalable, performant, and resilient
- Work closely with Infrastructure, Network, Security, Architecture, and Development teams to build highly performing, scalable, and secure Azure environments
- Define needs by documenting processes; includes research, planning and writing supporting documentation
Additional Functions: In addition to the essential functions listed above, the incumbent may perform the following additional functions.
- Participate in regulatory and compliance activities as necessary
Requirements
- Bachelor’s degree in Computer Science, Management Information Sciences or area of functional responsibility preferred, or equivalent years of industry work experience
- 5+ years in software or operations engineering
- 2+ years of DevOps and Site Reliability engineering or similar experience with cloud-native solutions
- Proven experience in DevOps culture and site reliability engineering focused on the customer, cross-functional autonomous teams, and continuous improvement
- DevOps experience with a cloud-native web application hosted in Microsoft Azure
- Familiarity with version control systems e.g., Git, SVN, CVS
- Extensive database and operating systems experience
- Experience in designing and implementing a continuous integration pipeline (CICD)
- Experience in monitoring infrastructure, application uptime, latency, and performance on large distributed systems
- Exhibit proficiency at troubleshooting various cloud and system related issues
- Demonstrable cross-functional knowledge with systems, storage, networking, security, and databases
-
Excellent verbal and written communication skills to convey monitoring insights and collaborate across technical and non-technical teams.
- Experience with cross-team collaboration. Partnering with DevOps/Platform Engineering, Production Support, and Architecture/Development teams to integrate monitoring solutions into existing applications, infrastructure, and automated pipelines
Preferred Experience:
-
Kubernetes Monitoring: Proven experience with observability solutions tailored to Kubernetes environments (AKS), including monitoring containerized services and workload
-
Monitoring Solutions: Experience with designing, deploying, and maintaining monitoring frameworks using Dynatrace, Azure Monitor, and Application Insights, ensuring comprehensive visibility across distributed systems
-
Alerting & Incident Response: Configured alerting mechanisms based on proactive symptom-based thresholds, enabling rapid resolution of performance and reliability issues
-
Performance Analysis: Experience analyzing telemetry and monitoring data to identify bottlenecks in application and system performance within the Kubernetes ecosystem
-
A passion for leveraging observability tools to drive operational improvements in cloud-native applications
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Is Software Engineering Over-Saturated?
Fully Remote Software Engineer Jobs
Where To Find Software Engineering Jobs