Staff Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Your core mandate: ensure Obsidian detects, diagnoses, and communicates system issues before customers are impacted-consistently and predictably. This is a hands-on technical role that involves architecting and leading the implementation of systems that handle real-world complexity, including upstream SaaS dependencies, sparse and noisy signals, and mission-critical enterprise workloads., * Reliability Strategy & Architecture - Define and lead long-term reliability strategy across services. Establish end-to-end system visibility frameworks and guide architecture for observability, detection, and resilience.
- Cross-Org Leadership - Partner across teams to embed reliability, standardize SLI/SLOs, and serve as a technical escalation expert.
- Detection & Observability - Build intelligent detection systems (anomaly detection, connector health models) and enable self-service observability.
- Incident Management - Define and evolve a tiered incident communication strategy, improve response practices, and lead postmortems to strengthen reliability and customer trust.
- Execution - Contribute hands-on to system design, monitoring, and debugging across distributed systems and data pipelines.
Requirements
- 5+ years in SRE, Production Engineering, or related roles
- 3+ years operating at a senior or technical leadership level (Staff or equivalent scope)
-
Deep expertise in:
- AWS and/or GCP
- Kubernetes and Helm
- Observability stacks (Prometheus, Grafana, or equivalent)
- CI/CD systems (GitLab CI/CD, ArgoCD, etc.)
Proven experience designing and scaling reliability systems for multi-tenant SaaS platforms
Strong debugging and systems thinking across distributed microservices and legacy systems
Demonstrated ability to lead initiatives that improve incident detection, response, and system resilience
Hands-on engineering approach with a track record of building-not just configuring-reliability systems, * Experience in B2B SaaS serving enterprise or financial customers
- Familiarity with third-party SaaS connector architectures and ingestion patterns
- Experience building anomaly detection or intelligent alerting systems
- Experience designing customer-facing status pages and incident communication frameworks
Benefits & conditions
Our competitive benefits packages are designed to support our employees’ well-being, both at work and at home. Our US-based employees enjoy competitive compensation with equity and 401k, comprehensive healthcare with dental and vision coverage, flexible paid time off and paid holiday time off, 12 weeks of new parent or family leave, and personal and professional development resources. For more details on our US benefits, or for information on our international benefits, please see here., Please note that the base pay range is a guideline and for candidates who receive an offer, the base pay will vary based on factors such as work location, as well as the knowledge, skills and experience of the candidate. In addition to a competitive base salary, this position is eligible for equity awards and may be eligible for sales commission or incentive compensation based on the role or function within the company.
Base Salary Range: £124,000 GBP - £141,000 GBP
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
Fully Remote Software Engineer Jobs
Dev Digest 120 - Apple and peers
The Best Software Developer Blogs to Read