> Markdown version of [/jobs/ext/2790860-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2790860-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SIte Reliability Engineer - **Company:** Arbor Education - **Location:** UK (Remote available) - **Contract:** Permanent contract - **Skills:** Clean Code Principles, Adobe InDesign, Agile Methodology, Databases, Relational Databases, DevOps, Nginx, Reliability Engineering, Site Reliability Engineering Practices, Prometheus, Software Engineering, Datadog, Test-Driven Development (TDD), System Availability, Containerization, AWS Aurora, Performance Monitor, Terraform, Domain Driven Design, Code Restructuring, Docker - **Published:** September 8, 2026 - **Apply:** https://startup.jobs/site-reliability-engineer-arbor-education-9951002 ## About the Role * Experience in performance monitoring and analysis * Capacity planning experience * Scripting and automation skills, with experience in relevant technologies. * Experience with Infrastructure as Code, in particular, Terraform * Understanding of relational database technologies and their cloud versions (e.g. AWS Aurora) * Experience with messaging and distributed asynchronous workloads * Experience with nginx or similar technologies * Familiarity with SRE processes. * Aware of DevOps principles like the 3 ways and 5 ideals. Bonus Skills * Experience with other database technologies and cloud platforms. * Past experience with Enterprise solutions running at scale * Familiarity with Kanban and Agile development processes * Experience with containerisation, for example Docker * Familiarity with software best practices such as Refactoring, Clean Code, Domain-Driven Design and Test-Driven Development. ## Description We are looking for an enthusiastic and proactive Site Reliability Engineer to join our SRE team and help us ensure we provide world-class resilience and performance across the platform. The remit and focus of the role is to advise on all aspects of site reliability including availability, scalability, observability and capacity planning. It's a broad and exciting role, so we're looking for someone up for a challenge - if you're an energetic and a collaborative Site Reliability Engineer, this is the role for you. Core responsibilities * Proactively monitor and analyse platform performance. * Collaborate with engineering teams to address performance bottlenecks and ensure scalability. * Assist engineering teams with implementing and reviewing SLOs * Continually improve observability through monitoring and alerting, and dashboards, using tools such as DataDog or Prometheus for example. * Work with other teams to ensure it is effective and provides full coverage. * Ensure the service is highly available and resilient * Champion best practices in design for high availability * Devise runbooks and run game sessions to test our DR plan, H/A and backups * Conduct assessments of capacity and plan for scaling to meet current and future business needs. * Work closely with the Head of Platform Engineering and Head of SRE to strategize and implement scalable solutions. * Work closely with the Platform team, feature teams and, 2nd line support and other stakeholders to ensure a good level of service is provided for our customers and embed SRE practices. * Key player in the response and troubleshooting of incidents, ensuring rapid resolution and minimising downtime. * Participate in blameless postmortems to identify root cause and corrective actions * Develop and maintain playbooks and documentation ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Post-Quantum Cryptography: Preparing for Q-Day](https://www.wearedevelopers.com/videos/100179-post-quantum-cryptography-preparing-for-q-day) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)