> Markdown version of [/jobs/ext/2244604-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2244604-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** Roche - **Location:** Madrid, Spain - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Continuous Integration, Distributed Systems, Fault Tolerance, Python (Programming Language), Reliability Engineering, Software Engineering, Scripting, Grafana, Mttr, Kubernetes, Information Technology, Terraform - **Published:** August 26, 2026 - **Apply:** https://www.buscojobs.com.es/senior-site-reliability-engineer-en-madrid-ID-368872114 ## About the Role Join Roche to contribute to a mission that prioritizes patient access and quality of care.Compensaciones / Beneficios* Define and implement SLIs, SLOs, and error budgets with product and engineering teams* Conduct reliability reviews for new and existing services* Design scalable, fault-tolerant architectures in AWS and Azure environments* Lead capacity planning, performance and cost optimization initiatives* Improve system resilience through automation and self-healing patterns* Drive organizational observability maturity (metrics, logs, traces, alert quality)* Perform complex root cause analysis and drive rapid mitigation* Participate in blameless postmortems and follow-through* Improve MTTR, reduce incident frequency, and elevate production standards* Collaborate with engineering teams to enable timely resolutions* Handle requests and incidents, create and maintain runbooks* Participation in a structured 24*7 on-call rotation* Reduce operational toil through tooling and automation (Python or similar)* Improve CI/CD reliability and deployment safety mechanisms* Build and maintain infrastructure-as-code (Terraform or equivalent)* Enhance Kubernetes platform reliability (EKS, AKS, or similar)* Partner with business, engineering, security, and cloud teams to embed reliability early in the software development life cycle* Mentor mid-level engineers and help shape SRE best practices* Championing a culture of ownership, accountability, and continuous improvement xqbhyrx Responsabilidades* Bachelor in Computer Science or related field or equivalent experience* Experience in site reliability engineering or software engineering with on-call experience* Strong AWS and/or Azure experience with cloud resources (Kubernetes, EKS, AKS, GKE)* Proficiency with observability tools* Hands-on incident management experience and tooling* Scripting skills for automation (e.g., Python)* Proven troubleshooting in cloud and distributed systems* Strong communication, teamwork, and documentation skills* Commitment to diversity and inclusive collaboration* Fluent English communicationRequisitos principales* ## Description Experteer Overview Descubra si esta oportunidad es adecuada para usted leyendo toda la información que sigue a continuación.In this role you will help design, build, and scale reliable distributed systems that power healthcare innovation.You will work closely with development teams to improve uptime, efficiency, and deployment processes while reducing operational toil.You'll lead incident management and drive improvements across monitoring, automation, and runbooks.This position offers the chance to influence system design and enable faster, safer delivery of services at a global scale.Join Roche to contribute to a mission that prioritizes patient access and quality of care.Compensaciones / Beneficios* Define and implement SLIs, SLOs, and error budgets with product and engineering teams* Conduct reliability reviews for new and existing services* Design scalable, fault-tolerant architectures in AWS and Azure environments* Lead capacity planning, performance and cost optimization initiatives* Improve system resilience through automation and self-healing patterns* Drive organizational observability maturity (metrics, logs, traces, alert quality)* Perform complex root cause analysis and drive rapid mitigation* Participate in blameless postmortems and follow-through* Improve MTTR, reduce incident frequency, and elevate production standards* Collaborate with engineering teams to enable timely resolutions* Handle requests and incidents, create and maintain runbooks* Participation in a structured 24*7 on-call rotation* Reduce operational toil through tooling and automation (Python or similar)* Improve CI/CD reliability and deployment safety mechanisms* Build and maintain infrastructure-as-code (Terraform or equivalent)* Enhance Kubernetes platform reliability (EKS, AKS, or similar)* Partner with business, engineering, security, and cloud teams to embed reliability early in the software development life cycle* Mentor mid-level engineers and help shape SRE best practices* Championing a culture of ownership, accountability, and continuous improvement xqbhyrx Responsabilidades* Bachelor in Computer Science or related field or equivalent experience* Experience in site reliability engineering or software engineering with on-call experience* Strong AWS and/or Azure experience with cloud resources (Kubernetes, EKS, AKS, GKE)* Proficiency with observability tools* Hands-on incident management experience and tooling* Scripting skills for automation (e.g., Python)* Proven troubleshooting in cloud and distributed systems* Strong communication, teamwork, and documentation skills* Commitment to diversity and inclusive collaboration* Fluent English communicationRequisitos principales* ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [JavaScript? No. Java Scripts! - Scripting with Java](https://www.wearedevelopers.com/videos/2094-javascript-no-java-scripts-scripting-with-java) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Intermediate Bitcoin Script](https://www.wearedevelopers.com/videos/25-intermediate-bitcoin-script) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers)