Site Reliability Engineer
Role details
Job location
Tech stack
Job description
At Allwyn, the Site Reliability Engineer supports the reliability and performance of digital services by operating production systems, building automation, and improving observability.
You will work closely with Senior SREs and engineering teams to ensure services remain stable, scalable, and well-instrumented across both steady-state and high-demand events.
What you'll be doing
Objectives of the role
- Maintain reliable production services across digital platforms
- Improve monitoring, alerting, and observability coverage
- Reduce operational toil through automation
- Support incident response and continuous improvement
- Contribute to performance and scaling of services
Production operations
-
Participate in 1-in-4 on-call rotation:
-
Respond to incidents
-
Support out-of-hours diagnosis
-
Monitor system health across:
-
Web and mobile platforms
-
Games platforms
-
Player services
Troubleshoot issues across application, infrastructure, and network layers
Incident response & improvement
- Support incident triage and resolution
- Participate in post-incident reviews and implement remediation actions
- Maintain and improve runbooks and operational documentation
Observability
-
Implement and maintain monitoring using:
-
Splunk
-
CloudWatch
-
Grafana
Improve:
- Logging quality
- Metrics coverage
- Alerting accuracy
Contribute to linking system performance to user experience signals
Automation & engineering
-
Develop scripts and tooling to:
-
Reduce manual tasks
-
Improve repeatability and reliability
Contribute to infrastructure management using TerraformSupport deployment processes and CI/CD improvements
Platform & cloud
- Work with AWS services (primarily ECS, with exposure to EKS/Kubernetes)
- Support scalability and availability improvements
- Assist in performance tuning and capacity planning
Collaboration
-
Work closely with engineers to:
-
Improve service reliability
-
Support releases and production readiness
Contribute to adoption of SRE practices within teams
Requirements
-
Experience in cloud environments (AWS preferred)
-
Working knowledge of:
-
Containers (ECS; exposure to Kubernetes is a plus)
-
Infrastructure as Code (Terraform)
Basic programming/scripting capability (Python, Bash, etc.)
Troubleshooting & operations
- Ability to diagnose issues in distributed systems
- Familiarity with monitoring and logging tools
- Understanding of Linux systems and networking fundamentals
SRE fundamentals
-
Understanding of:
-
Monitoring and alerting concepts
-
Reliability principles
-
Incident response processes
Willingness to be part of an on-call rotation
Desirable Experience
-
Exposure to:
-
SLOs / SLIs / error budgets
-
CI/CD pipelines
-
Performance and load testing
Familiarity with observability platforms used in modern digital estatesExperience in customer-facing, high-availability systems
Benefits & conditions
-
Company Bonus Scheme
-
Matched pension contributions up to 8.5%
-
26 days annual leave + 2 Life Days (and bank holidays)
-
Single Private Health Cover
-
Complimentary Private Medical
-
Income Protection
-
Flexible Benefits - EV Scheme, Money Coach, Will Writing, Mortgage Advice, Dental and Eye Care Schemes.
-
Enhanced Family Leave (Maternity, Paternity, Adoption)
-
Wellness Allowance £500
-
Employee Assistance Programme
-
Discounted Health Assessments
-
Volunteering Days
-
Matched Funding
We are a Disability Confident Leader which means we've taken proactive steps to ensure our workplace is accessible and inclusive for disabled and neurodivergent colleagues and candidates. As part of this we offer an interview to disabled applicants who meet the essential requirements of the job.