> Markdown version of [/jobs/ext/1379268-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1379268-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** Runware - **Location:** UK (Remote available) - **Experience:** Expert - **Salary:** £71,720.0 - **Contract:** Permanent contract - **Skills:** PHP (Programming Language), Application Programming Interfaces (APIs), Databases, Software Debugging, DevOps, Distributed Systems, Python (Programming Language), MySQL, Operational Databases, Queueing Systems, RabbitMQ, Redis, Reliability Engineering, Load Balancing, Kubernetes, Deployment Automation, Bare Metal, Vertica, Dynatrace - **Published:** July 22, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5807980864 ## About the Role * Have strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar role * Have a strong understanding of distributed systems and are comfortable debugging across applications, databases, queues, containers, networking and infrastructure * Have experience designing and operating observability systems using metrics, logs and distributed tracing * Understand SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing operational toil * Have experience with Kubernetes, containers, IaC and automated deployment practices, alongside the ability to write software and automation using languages such as Python, Go or PHP * Take strong ownership of production problems and are comfortable participating in an engineering on-call rotation, taking issues from initial investigation through to long-term remediation Bonus * Experience operating high-throughput or low-latency APIs and distributed systems * Experience with bare-metal infrastructure, GPU environments or AI and ML workloads * Experience with RabbitMQ or other distributed messaging and queueing systems * Experience operating MySQL, Redis, ClickHouse or similar production data systems * Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments * Experience building automated scaling, capacity management or self-healing systems ## Description As a Site Reliability Engineer at Runware, you will help ensure these systems remain reliable, performant and resilient as we scale. This is a highly technical, hands-on role working across software, infrastructure and production operations to improve observability, reduce incidents, eliminate operational toil and build lasting improvements across complex distributed systems. What you'll do * Own and improve the reliability, availability and performance of critical production services across the Runware platform * Define and evolve our reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards * Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in our engineering on-call rotation * Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements * Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience * Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements as the platform grows ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [MySQL Protocol Features You Should Be Aware Of](https://www.wearedevelopers.com/videos/100267-mysql-protocol-features-you-should-be-aware-of) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)