Principal Site Reliability Engineer (we have office locations in Cambridge, Leeds and London)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+8 more
Job description
As Principal Site Reliability Engineer, you will help establish and grow Genomics England’s organisation-wide SRE capability.
You will be responsible for:
- Building and leading a small team of Site Reliability Engineers (initially 3 people)
- Identifying the highest-value opportunities to improve reliability across our platforms
- Defining standards, patterns and tooling that help engineering teams to build and operate reliable services
- Developing services and capabilities that reduce operational toil and improve engineering effectiveness
- Partnering with product and engineering teams to embed SRE principles into their ways of working
- Influencing engineering strategy, governance and technical direction across the organisation
- Helping to build a culture of reliability, continuous improvement and operational resilience
The SRE team will work alongside product squads, helping them adopt reliability practices, improve operations and solve complex technical challenges, rather than operating services on their behalf.
In your first year you will establish the initial SRE team, identify the highest-priority reliability challenges across Genomics England, and define the standards, services and practices that will underpin our long-term approach to reliability.
The Principal SRE reports to the Director of Engineering within the Technology and Product Directorate.
About the Tech Stack
The new SRE team will support squads that run a variety of services: most of these are either user-facing web applications (React), backend APIs (Python), bioinformatics pipelines (NextFlow), or data ETL workflows (Prefect, Dremio). These services increasingly run in AWS, though there is still a significant on-premise presence, and they run in a mixture of compute environments, from ECS/Fargate to HPC clusters to (occasionally) Kubernetes.
Within the SDLC we have a standard toolchain which includes Terraform for infrastructure-as-code, GitLab for source code and CI/CD, Artifactory for software artefacts, and DataDog for observability. We are working to become interoperable with the wider NHS via open standards like FHIR and GA4GH APIs and increasingly aiming to integrate with their own API Management platform.
Requirements
You are a hands-on engineering leader with deep experience of SRE and platform engineering. You combine strategic thinking with the ability to identify and prioritise the areas where reliability improvements will have the greatest impact.
You are a problem-solver who identifies risks, issues, gaps, and dependencies and brings people together to find solutions. You do this through your supportive, empathetic and collaborative behaviours, acting as a coach, mentor, guide or constructive questioner as the situation demands. You are a great communicator, comfortable not just leading your own team, but also engaging across the engineering community and with non-technical stakeholders.
Essential Skills and Experience
While we recognise the value of relevant qualifications or certifications, we are primarily interested in your real-world experience:
- Comprehensive knowledge of SRE principles and practices with significant experience of applying these to real-world situations
- Excellent software engineering skills especially in the context of release automation and other toil-eliminating activities (Python preferred, polyglot ideal)
- Strong understanding of how architecture and other factors contribute to the overall resilience of systems
- Extensive experience of platform engineering across CI/CD, Infrastructure as Code, operational monitoring and alerting, backup and recovery etc.
- Experience in at least one major public cloud (AWS preferred but not essential)
- Demonstrable ability to lead teams, develop people and coordinate work towards shared outcomes
- Strong interpersonal skills with a temperament that builds trust and connection within and across squads through open, honest communication
- Comfortable engaging responsively with teams both remotely and in person when required
- Ability to navigate rapidly to effective solutions through engaged and inclusive listening, clarity of thought, clear documentation, and succinct presentation
Desirable Experience
These skills are not essential but if you have either of them, they may prove to be useful:
- Background in healthcare or bioinformatics
- Experience in regulated environments
About the company
If you’re an experienced Site Reliability Engineer leader, who thrives on working collaboratively to mature engineering practices, we’d love to hear from you. Join us at Genomics England and make a meaningful impact in the world of genomics.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
Data Engineer Salary UK
Software Engineer Salary London
Is Software Engineering Over-Saturated?