Site Reliability Engineer (SRE)

Xpertise Recruitment
West Drayton, UK
10 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Amazon Web Services Cloud Computing Continuous Integration Reliability Engineering Site Reliability Engineering Practices Ansible Datadog Cloudformation Containerization Kubernetes Information Technology Terraform
+3 more
Splunk Devsecops Docker

Job description

Overview

Site Reliability Engineer (SRE) Location: London

I am looking for a number of SREs for a large-scale digital organisation in the middle of a major engineering modernisation journey. This is not a BAU support role, this is a chance to help define what “good” looks like as SRE is brought fully in-house for the first time.

You’ll work across high-impact platforms (web/mobile, payments, CRM, operations, cloud) and play a key role in shifting the organisation away from ticket-driven support and towards proactive, automated, AWS-first, engineering-led reliability.

Responsibilities

  • Embed SRE principles to improve availability, reliability, performance and incident response
  • Modernise legacy support by introducing automation, observability, shift-left practices and CI/CD
  • Work across multiple domains (web/mobile, payments, CRM, cloud infrastructure, airline systems)
  • Partner with vendors and internal engineering teams; influence technical and financial decisions
  • Define and drive SLOs/SLIs, service health metrics and standards
  • Troubleshoot, monitor and improve systems using tooling such as Datadog, Splunk etc.
  • Contribute to IaC, containerisation and cloud-native adoption (AWS, Terraform, Docker/K8s)
  • Mentor and support engineers as SRE ways of working are introduced across the organisation

Experience

  • 3-5 years in an SRE or closely related reliability/DevSecOps discipline
  • Strong knowledge of SRE practices: monitoring, observability, incident response, automation
  • Hands-on with AWS and infrastructure-as-code (Terraform, Ansible or CloudFormation)
  • Experience with CI/CD pipelines and container platforms (Docker / Kubernetes)
  • Comfortable working with vendors, suppliers and internal product/engineering teams
  • Able to communicate clearly and influence engineering culture, not just build solutions
  • Exposure to large, complex or vendor-heavy environments is desirable
  • Hands-on role leadership through influence, not a line-manager post

Interested in learning more? Please get in touch with Benjamin Applewhaite to discuss the role in confidence.

Seniority level

  • Mid-Senior level

Employment type

  • Full-time

Job function

  • Information Technology

Industries

  • Technology, Information and Internet

Requirements

  • 3-5 years in an SRE or closely related reliability/DevSecOps discipline
  • Strong knowledge of SRE practices: monitoring, observability, incident response, automation
  • Hands-on with AWS and infrastructure-as-code (Terraform, Ansible or CloudFormation)
  • Experience with CI/CD pipelines and container platforms (Docker / Kubernetes)
  • Comfortable working with vendors, suppliers and internal product/engineering teams
  • Able to communicate clearly and influence engineering culture, not just build solutions
  • Exposure to large, complex or vendor-heavy environments is desirable
  • Hands-on role leadership through influence, not a line-manager post

Interested in learning more? Please get in touch with Benjamin Applewhaite to discuss the role in confidence.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

3:19 min

Executing complex workflows using Ansible Automation Platform

Goetz Rieger Goetz Rieger · World Congress 2025

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all