Senior Site Reliability Engineer - London

Spectrum IT Recruitment
London, UK
1 day ago
Apply on www.adzuna.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
£60,000.0 - £65,000.0
Working hours
Shift work

Tech stack

Artificial Intelligence Amazon Web Services Bash Shell Cloud Computing Linux DevOps Domain Name System (DNS) Monitoring of Systems Python (Programming Language) Linux System Administration Networking Basics Reliability Engineering
+11 more
Prometheus TCP/IP Datadog Load Balancing Cloud Platform System Grafana Kubernetes Cloudwatch Terraform Splunk Docker

Job description

  • Monitor and maintain highly available production platforms running in AWS
  • Respond to and manage production incidents across a 24/7 service
  • Investigate complex technical issues and restore services quickly and effectively
  • Develop automation to reduce manual operational tasks and improve platform resilience
  • Build and improve monitoring, alerting, and observability across cloud environments
  • Work alongside Software, Platform, Cloud, and Security Engineers to improve reliability and operational excellence
  • Contribute to post-incident reviews and drive continuous service improvements
  • Support containerised workloads using Kubernetes and Docker

Technologies:

  • AI
  • AWS
  • Bash
  • Cloud
  • CloudWatch
  • Datadog
  • Docker
  • Grafana
  • Support
  • Kubernetes
  • Linux
  • Load Balancing
  • Prometheus
  • Python
  • Security
  • Splunk
  • TCP/IP
  • Terraform
  • DevOps

Requirements

  • Experience in a Site Reliability Engineering, Production Engineering, Cloud Operations, or NOC environment
  • Exposure to Linux systems administration
  • Exposure to AWS cloud infrastructure
  • Exposure to Kubernetes and Docker
  • Exposure to production support and incident management
  • Exposure to Python, Bash, or Go scripting
  • Exposure to monitoring and observability platforms such as Grafana, Prometheus, Datadog, Splunk, or CloudWatch
  • Exposure to networking fundamentals including DNS, TCP/IP, and load balancing
  • A passion for automation, continuous improvement, and operational excellence
  • Experience with Infrastructure as Code such as Terraform would be beneficial
  • Experience with SRE principles such as SLIs and SLOs would be beneficial
  • Experience in regulated environments would be beneficial

Benefits & conditions

We are a global leader in AI-powered customer experience and cloud technology, and we are expanding our engineering teams following the award of a major government programme. We are building and supporting highly secure, cloud-native platforms that deliver sensitive communication services. This is a fully remote role in the UK on a 24/7 shift pattern with a 28-day rota including days and nights. We offer a competitive salary, bonus, and excellent benefits, and you will join an engineering-led organisation where reliability, automation, and continuous improvement are central to our platform.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · World Congress 2022

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all