> Markdown version of [/jobs/ext/1475043-senior-lead-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1475043-senior-lead-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior / Lead Site Reliability Engineer - **Company:** Allwyn UK - **Location:** Watford, UK - **Experience:** Expert - **Salary:** £77,970.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Computer Programming, Continuous Integration, Distributed Systems, Python (Programming Language), Log Analysis, Reliability Engineering, Site Reliability Engineering Practices, Scripting, Cloud Platform System, System Availability, Grafana, Mttr, Technical Debt, Containerization, Kubernetes, Low Latency, Performance Monitor, Cloudwatch, Terraform, Splunk, Stream Analytics - **Published:** July 29, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5818525886 ## About the Role * Strong experience in cloud environments, ideally AWS (ECS, with exposure or experience in EKS/Kubernetes) * Hands-on experience with: * Terraform (Infrastructure as Code) * Observability stacks (Splunk, CloudWatch, Grafana) Strong programming skills (Python, Go, or similar) SRE practices * Proven experience implementing: * SLOs, SLIs, error budgets * Incident management frameworks * Observability strategies Strong experience in distributed systems troubleshooting Operational * Experience in on-call production environments * Demonstrated leadership during high-severity incidents Desirable Experience * Experience migrating container platforms (ECS * EKS/Kubernetes) * Experience supporting high-scale consumer platforms * Familiarity with real-time analytics / customer experience tooling (e.g., Quantum Metric) * Experience in regulated or high-availability environments ## Description At Allwyn, the Senior/Lead Site Reliability Engineer is responsible for technical leadership of reliability engineering across the digital estate, ensuring high availability, performance, and resilience of customer-facing systems during both normal operation and peak lottery events. The role combines hands-on engineering, incident leadership, and ownership of the SRE improvement backlog and reporting, working across platform, product, and operational teams. Objectives of the role * Own reliability outcomes across services using SLOs, SLIs, and error budgets * Improve availability, latency, and scalability across Instant-Win and Draw-based platforms * Lead incident response and operational readiness, including peak jackpot events * Drive automation and platform maturity, reducing manual operational effort * Establish clear reporting on reliability, incidents, and service health trends * Supporting with thought-leadership and developing long-term roadmap What you'll be doing Reliability engineering & technical leadership * Define and govern SLOs / SLIs / error budgets across critical services * Lead reliability design reviews across: * Web & mobile platforms * Player Identity & Protection systems * CMS and Geolocation services Drive architecture improvements for resilience (failover, degradation, scaling patterns) Production operations & incident leadership * Act as incident commander for major incidents and high-severity events * Lead 1-in-4 on-call rotation, covering: * Out-of-hours incident diagnosis * Peak jackpot proactive monitoring Own end-to-end incident lifecycle: * Detection * triage * resolution * post-incident review Ensure blameless post-mortems with clear remediation ownership Observability & service insight * Define and evolve observability strategy using: * Splunk (log analytics) * CloudWatch (AWS telemetry) * Grafana (metrics visualisation) * Quantum Metric (user behaviour insight) Standardise: * Alerting quality and signal-to-noise ratio * Dashboards aligned to SLOs and customer impact * Drive correlation between technical signals and user experience Automation & platform engineering * Lead automation of operational processes using Terraform and scripting * Improve deployment and release safety (CI/CD, progressive delivery patterns) * Key contributor to transition strategy for ECS * EKS (Kubernetes adoption) Reduce operational toil through tooling, self-healing, observability and platform improvementsEmpower Level-1 operational teams with safe, controlled access to the tools they need to operate autonomously Capacity & performance engineering * Own capacity planning for: * High-concurrency draw events * Traffic spikes during jackpots Lead performance optimisation: * Latency reduction * Throughput scaling * Cost efficiency (AWS utilisation and associated log costs, observability license consumption) Backlog ownership & reporting * Own and prioritise the SRE backlog, balancing: * Reliability improvements * Technical debt * Automation opportunities to reduce/offload toil Produce structured reporting covering: * SLO performance * Incident trends and MTTR * Platform health and risk areas Provide clear updates to engineering leadership and business stakeholders Collaboration & culture * Embed SRE practices across engineering teams * Mentor engineers and SREs on: * Reliability engineering * Observability * Incident management Promote a culture of automation, measurement, and continuous improvementSupporting with thought-leadership and developing long-term SRE roadmap for Digital Operations ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fullstack Developer Salary UK](https://www.wearedevelopers.com/magazine/251-fullstack-developer-salary-uk) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk)