> Markdown version of [/jobs/ext/2938817-senior-site-reliability-engineer-roche](https://www.wearedevelopers.com/jobs/ext/2938817-senior-site-reliability-engineer-roche). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - Roche - **Company:** Roche - **Location:** Barcelona, Spain - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Cloud Computing, Continuous Integration, Distributed Systems, Fault Tolerance, Python (Programming Language), Reliability Engineering, Software Engineering, Scripting, Grafana, Software Troubleshooting, Kubernetes, Information Technology, Terraform - **Published:** September 16, 2026 - **Apply:** https://www.buscojobs.com.es/senior-site-reliability-engineer-roche-en-barcelona-ID-372136410 ## About the Role Requisitos principalesBachelor's degree in computer science, engineering, or related field or equivalent experience Production on-call experience in SRE or software engineering Experience with AWS and/or Azure (Kubernetes, EKS, AKS, GKE) Proficiency with observability tools Hands-on incident management tooling experience Scripting skills for automation (e.g., Python) Strong troubleshooting in cloud and distributed systems Excellent communication, teamwork, and documentation skills English proficiency Diversity and inclusion mindset strong communication teamwork proactive and self-motivated AWS and/or Azure cloud platforms Kubernetes (EKS/AKS/GKE) ## Description OverviewJoin Roche as a Site Reliability Engineer to design, build, and scale reliable distributed systems that power healthcare innovation.You will drive reliability, automation, and operational excellence, collaborating with product teams to improve uptime and system performance.Expect on-call rotation, blameless postmortems, and a focus on reducing toil through engineering solutions.This role offers a chance to shape the backbone of IT platforms at a company committed to global health impact.ResponsabilidadesDefine and implement SLIs, SLOs, and error budgets with product and engineering teamsConduct reliability reviews for new and existing servicesDesign scalable, fault-tolerant architectures in AWS and AzureLead capacity planning, performance and cost optimizationImprove system resilience through automation and self-healing patternsDrive observability maturity (metrics, logs, traces, alert quality)Incident management and continuous improvement through root cause analyses and postmortemsHandle requests and incidents, maintain runbooksParticipate in a 24x7 on-call rotationAutomation & platform engineering: reduce toil through tooling (Python or similar), improve CI/CD reliability, IaC with Terraform, enhance Kubernetes platforms (EKS/AKS/GKE)Cross-functional leadership: collaborate with business, security, and cloud teams; mentor engineers; promote ownership and continuous improvementRequisitos principalesBachelor's degree in computer science, engineering, or related field or equivalent experienceProduction on-call experience in SRE or software engineeringExperience with AWS and/or Azure (Kubernetes, EKS, AKS, GKE)Proficiency with observability toolsHands-on incident management tooling experienceScripting skills for automation (e.g., Python)Strong troubleshooting in cloud and distributed systemsExcellent communication, teamwork, and documentation skillsEnglish proficiencyDiversity and inclusion mindsetstrong communicationteamworkproactive and self-motivatedAWS and/or Azure cloud platformsKubernetes (EKS/AKS/GKE)Terraform or Infrastructure as Code ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [JavaScript? No. Java Scripts! - Scripting with Java](https://www.wearedevelopers.com/videos/2094-javascript-no-java-scripts-scripting-with-java) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries)