Senior Site Reliability Engineer - Roche

Roche
Barcelona, Spain
5 days ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English

Tech stack

Amazon Web Services Microsoft Azure Cloud Computing Continuous Integration Distributed Systems Fault Tolerance Python (Programming Language) Reliability Engineering Software Engineering Scripting Grafana Software Troubleshooting
+3 more
Kubernetes Information Technology Terraform

Job description

OverviewJoin Roche as a Site Reliability Engineer to design, build, and scale reliable distributed systems that power healthcare innovation.You will drive reliability, automation, and operational excellence, collaborating with product teams to improve uptime and system performance.Expect on-call rotation, blameless postmortems, and a focus on reducing toil through engineering solutions.This role offers a chance to shape the backbone of IT platforms at a company committed to global health impact.ResponsabilidadesDefine and implement SLIs, SLOs, and error budgets with product and engineering teamsConduct reliability reviews for new and existing servicesDesign scalable, fault-tolerant architectures in AWS and AzureLead capacity planning, performance and cost optimizationImprove system resilience through automation and self-healing patternsDrive observability maturity (metrics, logs, traces, alert quality)Incident management and continuous improvement through root cause analyses and postmortemsHandle requests and incidents, maintain runbooksParticipate in a 24x7 on-call rotationAutomation & platform engineering: reduce toil through tooling (Python or similar), improve CI/CD reliability, IaC with Terraform, enhance Kubernetes platforms (EKS/AKS/GKE)Cross-functional leadership: collaborate with business, security, and cloud teams; mentor engineers; promote ownership and continuous improvementRequisitos principalesBachelor’s degree in computer science, engineering, or related field or equivalent experienceProduction on-call experience in SRE or software engineeringExperience with AWS and/or Azure (Kubernetes, EKS, AKS, GKE)Proficiency with observability toolsHands-on incident management tooling experienceScripting skills for automation (e.g., Python)Strong troubleshooting in cloud and distributed systemsExcellent communication, teamwork, and documentation skillsEnglish proficiencyDiversity and inclusion mindsetstrong communicationteamworkproactive and self-motivatedAWS and/or Azure cloud platformsKubernetes (EKS/AKS/GKE)Terraform or Infrastructure as Code

Requirements

Requisitos principalesBachelor’s degree in computer science, engineering, or related field or equivalent experience Production on-call experience in SRE or software engineering Experience with AWS and/or Azure (Kubernetes, EKS, AKS, GKE) Proficiency with observability tools Hands-on incident management tooling experience Scripting skills for automation (e.g., Python) Strong troubleshooting in cloud and distributed systems Excellent communication, teamwork, and documentation skills English proficiency Diversity and inclusion mindset strong communication teamwork proactive and self-motivated AWS and/or Azure cloud platforms Kubernetes (EKS/AKS/GKE)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:04 min

Introduction to Bitcoin script parsing tools

Steve Shadders · LIVE

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

Videos

See all

Related articles

See all