Head of Site Reliability Engineering (SRE)

Computershare
Bristol, UK
22 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Shift work
Job source

Tech stack

Microsoft Azure Bash Shell Continuous Integration Python (Programming Language) Windows PowerShell Reliability Engineering Site Reliability Engineering Practices Ansible Prometheus Software Engineering Scripting Grafana
+5 more
Gitlab-ci Terraform Splunk Dynatrace Jenkins

Job description

In this position, you’ll be based in the Bristol or Edinburgh office for a minimum of three days a week, with the flexibility to work from home for some of your working week. Find out more about our flexible work culture at computershare.com/flex.

We give you a world of potential

Computershare have an up-and-coming opportunity for a Head of Site Reliability Engineering (SRE) to join our global technology team at a time when transforming our organisations toward an SRE operating model is a key focus.

Reporting directly to the Global Head of Technology Operations, you will operate within Technology Services with a global mandate to establish and mature Site Reliability Engineering capabilities across the organisation. Partnering closely with Engineering, Infrastructure Operations, and Security to improve service reliability, resilience, and performance of critical platforms.

A role you will love

We are seeking an experienced and visionary Head of SRE to define, lead, and evolve our global reliability strategy. This is a senior leadership role responsible for driving operational excellence, service reliability, observability, automation, and continuous improvement across our technology landscape.

Working closely with Engineering, Infrastructure, Security, and Technology Operations teams, you will establish and embed modern SRE practices that enable highly reliable, scalable, and resilient services while fostering a culture of shared ownership and continuous learning.

Other key responsibilities:

  • Drive adoption of SRE principles (SLOs, error budgets, toil reduction).
  • Establish observability and monitoring standards.
  • Lead automation-first operations.
  • Improve incident and problem management maturity.
  • Partner with software and infrastructure engineering teams to embed reliability into the product lifecycle.Establish SRE governance, standards, and operating model.

Requirements

You’ll be an experienced SRE leader who combines deep technical expertise with the ability to build high-performing teams and drive operational transformation at scale.

Proven experience building, leading, and developing Site Reliability Engineering or Production Engineering teams, with a strong understanding of SRE principles including service level objectives (SLOs), service level indicators (SLIs), error budgets, and toil reduction.

Extensive experience driving automation initiatives, supported by strong scripting and development capabilities using technologies such as Python, PowerShell, Bash, Terraform and Ansible Automation Platform (AAP).

Some other key skills that you’ll have:

  • Robust knowledge of observability and monitoring practices, with hands-on experience implementing and managing platforms such as Dynatrace, Prometheus, Grafana, and Splunk.
  • Good understanding of CI/CD tooling and modern software delivery practices, including Jenkins, GitLab CI, and Azure DevOps.
  • A background spanning both software engineering and technology operations environments.
  • Professional certifications in cloud technologies, Site Reliability Engineering, platform engineering, or reliability engineering disciplines.
  • Passionate about reliability, resilience, automation, and continuous improvement.A strategic thinker who can balance long-term vision with operational delivery.

If you’re confident leader able to inspire teams, challenge traditional ways of working, and drive meaningful change, we’d love to hear from you.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on uk.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier · World Congress 2023

6:58 min

Building engineering communities and finding technical inspiration

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

Videos

See all

Related articles

See all