Senior Site Reliability Engineer - London
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+11 more
Job description
- Monitor and maintain highly available production platforms running in AWS
- Respond to and manage production incidents across a 24/7 service
- Investigate complex technical issues and restore services quickly and effectively
- Develop automation to reduce manual operational tasks and improve platform resilience
- Build and improve monitoring, alerting, and observability across cloud environments
- Work alongside Software, Platform, Cloud, and Security Engineers to improve reliability and operational excellence
- Contribute to post-incident reviews and drive continuous service improvements
- Support containerised workloads using Kubernetes and Docker
Technologies:
- AI
- AWS
- Bash
- Cloud
- CloudWatch
- Datadog
- Docker
- Grafana
- Support
- Kubernetes
- Linux
- Load Balancing
- Prometheus
- Python
- Security
- Splunk
- TCP/IP
- Terraform
- DevOps
Requirements
- Experience in a Site Reliability Engineering, Production Engineering, Cloud Operations, or NOC environment
- Exposure to Linux systems administration
- Exposure to AWS cloud infrastructure
- Exposure to Kubernetes and Docker
- Exposure to production support and incident management
- Exposure to Python, Bash, or Go scripting
- Exposure to monitoring and observability platforms such as Grafana, Prometheus, Datadog, Splunk, or CloudWatch
- Exposure to networking fundamentals including DNS, TCP/IP, and load balancing
- A passion for automation, continuous improvement, and operational excellence
- Experience with Infrastructure as Code such as Terraform would be beneficial
- Experience with SRE principles such as SLIs and SLOs would be beneficial
- Experience in regulated environments would be beneficial
Benefits & conditions
We are a global leader in AI-powered customer experience and cloud technology, and we are expanding our engineering teams following the award of a major government programme. We are building and supporting highly secure, cloud-native platforms that deliver sensitive communication services. This is a fully remote role in the UK on a 24/7 shift pattern with a 28-day rota including days and nights. We offer a competitive salary, bonus, and excellent benefits, and you will join an engineering-led organisation where reliability, automation, and continuous improvement are central to our platform.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Best Companies to work for in London: Top 25 Companies in 2023
Fully Remote Software Engineer Jobs
Is Software Engineering Over-Saturated?
Data Engineer Salary UK