SRE Technical Lead

83zero Ltd
Maidenhead, United Kingdom
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
£ 100K

Job location

Remote
Maidenhead, United Kingdom

Tech stack

Cloud Computing
Databases
Continuous Integration
DevOps
PostgreSQL
Openshift
Red Hat Enterprise Linux - RHEL
Reliability Engineering
Site Reliability Engineering Practices
Prometheus
Service Design
Datadog
Istio
Grafana
Multi-Cloud
SC Clearance
Kubernetes

Job description

  • Define and drive our SRE strategy, standards, SLAs, SLOs and error budgets.
  • Embed reliability engineering principles into our platform and service design.
  • Lead the adoption of core SRE practices including reliability reviews, operational readiness and toil reduction.
  • Drive automation across monitoring, incident response, recovery and remediation.
  • Govern reliability-focused Infrastructure as Code, CI/CD pipelines and operational tooling.
  • Identify and remove systemic causes of operational overhead while improving scalability, resilience and operability.
  • Act as the senior technical escalation point for major incidents and high-risk releases.
  • Lead blameless post-incident reviews and ensure measurable service improvements.
  • Define and oversee observability, monitoring and capacity management practices.
  • Ensure our SRE approaches align with security, governance and compliance requirements.
  • Mentor and coach senior engineers, helping to improve SRE maturity across engineering teams.

Technologies:

  • ArgoCD
  • CI/CD
  • Cloud
  • GitOps
  • Grafana
  • Helm
  • Istio
  • Kubernetes
  • Kustomize
  • OpenTelemetry
  • OpenShift
  • PostgreSQL
  • Prometheus
  • Security
  • DevOps, We are hiring an SRE Technical Lead to take ownership of the reliability, availability and operational excellence of critical platforms within complex, multi-vendor environments. This is a senior technical leadership role where you will act as the technical authority for Site Reliability Engineering, working closely with stakeholders, engineering teams and delivery partners to drive reliability across large-scale cloud platforms. The position is hybrid in the UK, with office, client site and home-based working. We offer a salary of up to £100,000, a 5% annual bonus, and the opportunity to lead reliability engineering across large-scale, business-critical platforms while working with modern cloud-native technologies and complex enterprise environments. Please note that active SC clearance and sole UK nationality are mandatory requirements for this position.

Requirements

  • Hold active SC (Security Check) clearance.
  • Be a sole UK national.
  • Have strong technical expertise gained within enterprise-scale environments.
  • Have deep knowledge of Kubernetes and OpenShift.
  • Have experience designing and supporting hybrid and multi-cloud platforms.
  • Have experience with service mesh technologies such as Istio.
  • Have strong hands-on experience with observability tooling including Prometheus, Grafana, Loki, Tempo and OpenTelemetry.
  • Have Infrastructure as Code and GitOps expertise using tools such as Helm, Kustomize, ArgoCD and Tekton.
  • Have experience building and improving CI/CD pipelines with a focus on reliability engineering.
  • Have familiarity with Red Hat ACM/ACS, Submariner networking and enterprise databases such as PostgreSQL.
  • Have a proven track record of providing technical leadership across complex, multi-vendor environments.

Apply for this position