Staff Site Reliability Engineer

Obsidian Security
Cheltenham, UK
6 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
£124,000.0 - £141,000.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Software as a Service Continuous Integration Distributed Systems Reliability Engineering Prometheus Grafana Gitlab-ci Kubernetes Legacy Systems

Job description

Your core mandate: ensure Obsidian detects, diagnoses, and communicates system issues before customers are impacted-consistently and predictably. This is a hands-on technical role that involves architecting and leading the implementation of systems that handle real-world complexity, including upstream SaaS dependencies, sparse and noisy signals, and mission-critical enterprise workloads., * Reliability Strategy & Architecture - Define and lead long-term reliability strategy across services. Establish end-to-end system visibility frameworks and guide architecture for observability, detection, and resilience.

  • Cross-Org Leadership - Partner across teams to embed reliability, standardize SLI/SLOs, and serve as a technical escalation expert.
  • Detection & Observability - Build intelligent detection systems (anomaly detection, connector health models) and enable self-service observability.
  • Incident Management - Define and evolve a tiered incident communication strategy, improve response practices, and lead postmortems to strengthen reliability and customer trust.
  • Execution - Contribute hands-on to system design, monitoring, and debugging across distributed systems and data pipelines.

Requirements

  • 5+ years in SRE, Production Engineering, or related roles
  • 3+ years operating at a senior or technical leadership level (Staff or equivalent scope)
  • Deep expertise in:

  • AWS and/or GCP
  • Kubernetes and Helm
  • Observability stacks (Prometheus, Grafana, or equivalent)
  • CI/CD systems (GitLab CI/CD, ArgoCD, etc.)

Proven experience designing and scaling reliability systems for multi-tenant SaaS platforms

Strong debugging and systems thinking across distributed microservices and legacy systems

Demonstrated ability to lead initiatives that improve incident detection, response, and system resilience

Hands-on engineering approach with a track record of building-not just configuring-reliability systems, * Experience in B2B SaaS serving enterprise or financial customers

  • Familiarity with third-party SaaS connector architectures and ingestion patterns
  • Experience building anomaly detection or intelligent alerting systems
  • Experience designing customer-facing status pages and incident communication frameworks

Benefits & conditions

Our competitive benefits packages are designed to support our employees’ well-being, both at work and at home. Our US-based employees enjoy competitive compensation with equity and 401k, comprehensive healthcare with dental and vision coverage, flexible paid time off and paid holiday time off, 12 weeks of new parent or family leave, and personal and professional development resources. For more details on our US benefits, or for information on our international benefits, please see here., Please note that the base pay range is a guideline and for candidates who receive an offer, the base pay will vary based on factors such as work location, as well as the knowledge, skills and experience of the candidate. In addition to a competitive base salary, this position is eligible for equity awards and may be eligible for sales commission or incentive compensation based on the role or function within the company.

Base Salary Range: £124,000 GBP - £141,000 GBP

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

1:05 min

Measuring system availability utilizing Prometheus and straightforward PromQL

Alexander Schwartz Alexander Schwartz · World Congress 2025

13:07 min

Configuring application observability with Micrometer and Prometheus

Aleksandr Kalikov · LIVE

4:47 min

Automating frontend performance metrics with Google Lighthouse

Miki Lombardi · JS Congress

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

Videos

See all

Related articles

See all