Principal Site Reliability Engineer (we have office locations in Cambridge, Leeds and London)

Genomics England
London, UK
3 days ago
Apply on www.adzuna.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
£80,964.0
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Application Release Automation Continuous Integration Extract Transform Load (ETL) Python (Programming Language) Systems Development Life Cycle Reliability Engineering Software Engineering Web Applications Datadog ReactJS
+8 more
Fast Healthcare Interoperability Resources Gitlab Kubernetes AWS Fargate Terraform Api Management Artifactory Web Api

Job description

As Principal Site Reliability Engineer, you will help establish and grow Genomics England’s organisation-wide SRE capability.

You will be responsible for:

  • Building and leading a small team of Site Reliability Engineers (initially 3 people)
  • Identifying the highest-value opportunities to improve reliability across our platforms
  • Defining standards, patterns and tooling that help engineering teams to build and operate reliable services
  • Developing services and capabilities that reduce operational toil and improve engineering effectiveness
  • Partnering with product and engineering teams to embed SRE principles into their ways of working
  • Influencing engineering strategy, governance and technical direction across the organisation
  • Helping to build a culture of reliability, continuous improvement and operational resilience

The SRE team will work alongside product squads, helping them adopt reliability practices, improve operations and solve complex technical challenges, rather than operating services on their behalf.

In your first year you will establish the initial SRE team, identify the highest-priority reliability challenges across Genomics England, and define the standards, services and practices that will underpin our long-term approach to reliability.

The Principal SRE reports to the Director of Engineering within the Technology and Product Directorate.

About the Tech Stack

The new SRE team will support squads that run a variety of services: most of these are either user-facing web applications (React), backend APIs (Python), bioinformatics pipelines (NextFlow), or data ETL workflows (Prefect, Dremio). These services increasingly run in AWS, though there is still a significant on-premise presence, and they run in a mixture of compute environments, from ECS/Fargate to HPC clusters to (occasionally) Kubernetes.

Within the SDLC we have a standard toolchain which includes Terraform for infrastructure-as-code, GitLab for source code and CI/CD, Artifactory for software artefacts, and DataDog for observability. We are working to become interoperable with the wider NHS via open standards like FHIR and GA4GH APIs and increasingly aiming to integrate with their own API Management platform.

Requirements

You are a hands-on engineering leader with deep experience of SRE and platform engineering. You combine strategic thinking with the ability to identify and prioritise the areas where reliability improvements will have the greatest impact.

You are a problem-solver who identifies risks, issues, gaps, and dependencies and brings people together to find solutions. You do this through your supportive, empathetic and collaborative behaviours, acting as a coach, mentor, guide or constructive questioner as the situation demands. You are a great communicator, comfortable not just leading your own team, but also engaging across the engineering community and with non-technical stakeholders.

Essential Skills and Experience

While we recognise the value of relevant qualifications or certifications, we are primarily interested in your real-world experience:

  • Comprehensive knowledge of SRE principles and practices with significant experience of applying these to real-world situations
  • Excellent software engineering skills especially in the context of release automation and other toil-eliminating activities (Python preferred, polyglot ideal)
  • Strong understanding of how architecture and other factors contribute to the overall resilience of systems
  • Extensive experience of platform engineering across CI/CD, Infrastructure as Code, operational monitoring and alerting, backup and recovery etc.
  • Experience in at least one major public cloud (AWS preferred but not essential)
  • Demonstrable ability to lead teams, develop people and coordinate work towards shared outcomes
  • Strong interpersonal skills with a temperament that builds trust and connection within and across squads through open, honest communication
  • Comfortable engaging responsively with teams both remotely and in person when required
  • Ability to navigate rapidly to effective solutions through engaged and inclusive listening, clarity of thought, clear documentation, and succinct presentation

Desirable Experience

These skills are not essential but if you have either of them, they may prove to be useful:

  • Background in healthcare or bioinformatics
  • Experience in regulated environments

About the company

If you’re an experienced Site Reliability Engineer leader, who thrives on working collaboratively to mature engineering practices, we’d love to hear from you. Join us at Genomics England and make a meaningful impact in the world of genomics.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

6:14 min

Structuring CI/CD pipelines with integrated security and quality checks

Christoph Ruggenthaler · LIVE

9:56 min

Expanding browser capabilities with modern web APIs

Ire Aderinokun · JS Congress

1:21 min

Exploring the target application for front end tests

Anna Mcdougall · JS Congress

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · World Congress 2021

4:54 min

Implementing geographic salary tiers for compensation equity and fairness

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

Videos

See all

Related articles

See all