> Markdown version of [/jobs/ext/2307830-staff-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2307830-staff-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Site Reliability Engineer - **Company:** Obsidian Security - **Location:** Manchester, UK - **Experience:** Experienced - **Salary:** £124,000.0 - £141,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Software as a Service, Continuous Integration, Distributed Systems, Reliability Engineering, Prometheus, Grafana, Gitlab-ci, Kubernetes, Legacy Systems - **Published:** August 30, 2026 - **Apply:** https://www.collegerecruiter.com/job/2815108764-staff-site-reliability-engineer ## About the Role * 5+ years in SRE, Production Engineering, or related roles * 3+ years operating at a senior or technical leadership level (Staff or equivalent scope) * Deep expertise in: * AWS and/or GCP * Kubernetes and Helm * Observability stacks (Prometheus, Grafana, or equivalent) * CI/CD systems (GitLab CI/CD, ArgoCD, etc.) Proven experience designing and scaling reliability systems for multi-tenant SaaS platforms Strong debugging and systems thinking across distributed microservices and legacy systems Demonstrated ability to lead initiatives that improve incident detection, response, and system resilience Hands-on engineering approach with a track record of building-not just configuring-reliability systems, * Experience in B2B SaaS serving enterprise or financial customers * Familiarity with third-party SaaS connector architectures and ingestion patterns * Experience building anomaly detection or intelligent alerting systems * Experience designing customer-facing status pages and incident communication frameworks ## Description Your core mandate: ensure Obsidian detects, diagnoses, and communicates system issues before customers are impacted-consistently and predictably. This is a hands-on technical role that involves architecting and leading the implementation of systems that handle real-world complexity, including upstream SaaS dependencies, sparse and noisy signals, and mission-critical enterprise workloads., * Reliability Strategy & Architecture - Define and lead long-term reliability strategy across services. Establish end-to-end system visibility frameworks and guide architecture for observability, detection, and resilience. * Cross-Org Leadership - Partner across teams to embed reliability, standardize SLI/SLOs, and serve as a technical escalation expert. * Detection & Observability - Build intelligent detection systems (anomaly detection, connector health models) and enable self-service observability. * Incident Management - Define and evolve a tiered incident communication strategy, improve response practices, and lead postmortems to strengthen reliability and customer trust. * Execution - Contribute hands-on to system design, monitoring, and debugging across distributed systems and data pipelines. ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Data binning and understanding histograms](https://www.wearedevelopers.com/videos/2086-data-binning-and-understanding-histograms) - [Keycloak case study: Making users happy with service level indicators and observability](https://www.wearedevelopers.com/videos/1599-keycloak-case-study-making-users-happy-with-service-level-indicators-and-observability) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)