Head of Site Reliability Engineering (SRE)

Computershare
Bristol, United Kingdom
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Shift work
Languages
English
Experience level
Senior

Job location

Remote
Bristol, United Kingdom

Tech stack

Azure
Bash
Continuous Integration
Python
Powershell
Reliability Engineering
Site Reliability Engineering Practices
Ansible
Prometheus
Software Engineering
Scripting (Bash/Python/Go/Ruby)
Grafana
Gitlab-ci
Terraform
Splunk
Dynatrace
Jenkins

Job description

In this position, you'll be based in the Bristol or Edinburgh office for a minimum of three days a week, with the flexibility to work from home for some of your working week. Find out more about our flexible work culture at computershare.com/flex.

We give you a world of potential

Computershare have an up-and-coming opportunity for a Head of Site Reliability Engineering (SRE) to join our global technology team at a time when transforming our organisations toward an SRE operating model is a key focus.

Reporting directly to the Global Head of Technology Operations, you will operate within Technology Services with a global mandate to establish and mature Site Reliability Engineering capabilities across the organisation. Partnering closely with Engineering, Infrastructure Operations, and Security to improve service reliability, resilience, and performance of critical platforms.

A role you will love

We are seeking an experienced and visionary Head of SRE to define, lead, and evolve our global reliability strategy. This is a senior leadership role responsible for driving operational excellence, service reliability, observability, automation, and continuous improvement across our technology landscape.

Working closely with Engineering, Infrastructure, Security, and Technology Operations teams, you will establish and embed modern SRE practices that enable highly reliable, scalable, and resilient services while fostering a culture of shared ownership and continuous learning.

Other key responsibilities:

  • Drive adoption of SRE principles (SLOs, error budgets, toil reduction).
  • Establish observability and monitoring standards.
  • Lead automation-first operations.
  • Improve incident and problem management maturity.
  • Partner with software and infrastructure engineering teams to embed reliability into the product lifecycle.Establish SRE governance, standards, and operating model.

Requirements

You'll be an experienced SRE leader who combines deep technical expertise with the ability to build high-performing teams and drive operational transformation at scale.

Proven experience building, leading, and developing Site Reliability Engineering or Production Engineering teams, with a strong understanding of SRE principles including service level objectives (SLOs), service level indicators (SLIs), error budgets, and toil reduction.

Extensive experience driving automation initiatives, supported by strong scripting and development capabilities using technologies such as Python, PowerShell, Bash, Terraform and Ansible Automation Platform (AAP).

Some other key skills that you'll have:

  • Robust knowledge of observability and monitoring practices, with hands-on experience implementing and managing platforms such as Dynatrace, Prometheus, Grafana, and Splunk.
  • Good understanding of CI/CD tooling and modern software delivery practices, including Jenkins, GitLab CI, and Azure DevOps.
  • A background spanning both software engineering and technology operations environments.
  • Professional certifications in cloud technologies, Site Reliability Engineering, platform engineering, or reliability engineering disciplines.
  • Passionate about reliability, resilience, automation, and continuous improvement.A strategic thinker who can balance long-term vision with operational delivery.

If you're confident leader able to inspire teams, challenge traditional ways of working, and drive meaningful change, we'd love to hear from you.

Apply for this position