Senior Site Reliability Engineer - Roche
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+3 more
Job description
OverviewJoin Roche as a Site Reliability Engineer to design, build, and scale reliable distributed systems that power healthcare innovation.You will drive reliability, automation, and operational excellence, collaborating with product teams to improve uptime and system performance.Expect on-call rotation, blameless postmortems, and a focus on reducing toil through engineering solutions.This role offers a chance to shape the backbone of IT platforms at a company committed to global health impact.ResponsabilidadesDefine and implement SLIs, SLOs, and error budgets with product and engineering teamsConduct reliability reviews for new and existing servicesDesign scalable, fault-tolerant architectures in AWS and AzureLead capacity planning, performance and cost optimizationImprove system resilience through automation and self-healing patternsDrive observability maturity (metrics, logs, traces, alert quality)Incident management and continuous improvement through root cause analyses and postmortemsHandle requests and incidents, maintain runbooksParticipate in a 24x7 on-call rotationAutomation & platform engineering: reduce toil through tooling (Python or similar), improve CI/CD reliability, IaC with Terraform, enhance Kubernetes platforms (EKS/AKS/GKE)Cross-functional leadership: collaborate with business, security, and cloud teams; mentor engineers; promote ownership and continuous improvementRequisitos principalesBachelor’s degree in computer science, engineering, or related field or equivalent experienceProduction on-call experience in SRE or software engineeringExperience with AWS and/or Azure (Kubernetes, EKS, AKS, GKE)Proficiency with observability toolsHands-on incident management tooling experienceScripting skills for automation (e.g., Python)Strong troubleshooting in cloud and distributed systemsExcellent communication, teamwork, and documentation skillsEnglish proficiencyDiversity and inclusion mindsetstrong communicationteamworkproactive and self-motivatedAWS and/or Azure cloud platformsKubernetes (EKS/AKS/GKE)Terraform or Infrastructure as Code
Requirements
Requisitos principalesBachelor’s degree in computer science, engineering, or related field or equivalent experience Production on-call experience in SRE or software engineering Experience with AWS and/or Azure (Kubernetes, EKS, AKS, GKE) Proficiency with observability tools Hands-on incident management tooling experience Scripting skills for automation (e.g., Python) Strong troubleshooting in cloud and distributed systems Excellent communication, teamwork, and documentation skills English proficiency Diversity and inclusion mindset strong communication teamwork proactive and self-motivated AWS and/or Azure cloud platforms Kubernetes (EKS/AKS/GKE)
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Where To Find Software Engineering Jobs
Find a Developer Job: 12 Best Job Sites For Developers
Is Software Engineering Over-Saturated?
The 12 Best Jobs for Software Engineers