Principal/Senior Site Reliability Engineer

Roche
Charing Cross, United Kingdom
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior

Job location

Charing Cross, United Kingdom

Tech stack

Airflow
Amazon Web Services (AWS)
Azure
Bash
Catalyst
Cloud Computing
Cloud Engineering
Information Engineering
Disaster Recovery
Python
Reliability Engineering
TensorFlow
Prometheus
Datadog
Data Logging
Pulumi
Scripting (Bash/Python/Go/Ruby)
Load Balancing
High Performance Computing
Autoscaling
Grafana
Spark
Cloudformation
Kubernetes
Information Technology
Machine Learning Operations
Terraform
Serverless Computing
Docker

Job description

Join the Computational Sciences Center of Excellence as a Senior Site Reliability Engineer, where the platforms you build accelerate the discovery of transformative medicines. You will work alongside talented engineers in the Data & Digital Catalyst organisation to design resilient, cloud-based systems for MLOps and HPC workloads at global scale. This is a role for someone who wants their engineering craft to have real impact on science and patients., * You architect Infrastructure as Code using Terraform, Pulumi, or CloudFormation to provision and manage cloud infrastructure for MLOps and HPC workloads across global regions.

  • You design for resilience building disaster recovery and failover plans with auto-scaling and load balancing to keep critical systems available worldwide.
  • You strengthen reliability through chaos engineering running experiments that validate systems and surface weaknesses before they become incidents.
  • You build deep observability with monitoring, logging, and alerting frameworks such as Prometheus, Grafana, Datadog, and ELK.
  • You provide technical leadership to a team of engineers, fostering collaboration, innovation, and continuous improvement.
  • You partner across teams to align infrastructure with ML and HPC needs and to advance operational maturity through SLAs, SLOs, SLIs, and error budgets.

Requirements

  • You bring deep expertise in Infrastructure as Code with proven success deploying Terraform, Pulumi, or CloudFormation in AWS, Azure, or GCP for MLOps and HPC workloads.
  • You understand cloud-native and on-prem architectures including autoscaling, serverless, and multi-region deployments, and you are hands-on with Docker, Kubernetes, and Kubeflow.
  • You are an expert in automation scripting confidently in Python, Bash, or Go, with a strong grasp of GPU-accelerated computing and HPC workload scaling.
  • You lead through influence communicating and mentoring with clarity, and solving complex problems with a methodical approach.
  • You hold a degree in Computer Science or a related technical field or bring equivalent experience in software and site reliability engineering.

Preferred:

  • Experience with distributed ML frameworks such as Horovod or TensorFlow Distributed.
  • Familiarity with data engineering pipelines such as Apache Airflow or Apache Spark.
  • Knowledge of chaos engineering tools and compliance frameworks such as GDPR, SOC 2, or ISO 27001.

About the company

A healthier future drives us to innovate. Together, more than 100'000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come. Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a global impact. Let's build a healthier future, together. The statements herein are intended to describe the general nature and level of work being performed by employees, and are not to be construed as an exhaustive list of responsibilities, duties, and skills required of personnel so classified. Furthermore, they do not establish a contract for employment and are subject to change at the discretion of Roche Products Ltd. At Roche Products we believe diversity drives innovation and we are committed to building a diverse and flexible working environment. All qualified applicants will receive consideration for employment without regard to race, religion or belief, sex, gender reassignment, sexual orientation, marriage and civil partnership, pregnancy and maternity, disability or age. We recognise the importance of flexible working and will review all applicants' requests with care. At Roche difference is valued and we are proud to be an equal opportunity employer where you are encouraged to bring your whole self to work.

Apply for this position