> Markdown version of [/jobs/ext/1316408-staff-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1316408-staff-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Site Reliability Engineer - **Company:** Obsidian Security, Inc. - **Location:** United States - **Experience:** Expert - **Salary:** $232,000.0 - $263,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Software as a Service, Continuous Integration, Software Debugging, DevOps, Distributed Systems, Prometheus, Google Cloud, Grafana, Gitlab-ci, Kubernetes, Data Pipelines, Legacy Systems - **Published:** July 17, 2026 - **Apply:** https://www.dice.com/job-detail/80f64cb7-efc4-495f-82c1-fba007e68c79 ## About the Role * 5+ years in SRE, Production Engineering, or related roles * 3+ years operating at a senior or technical leadership level (Staff or equivalent scope) * Deep expertise in: * AWS and/or Google Cloud Platform * Kubernetes and Helm * Observability stacks (Prometheus, Grafana, or equivalent) * CI/CD systems (GitLab CI/CD, ArgoCD, etc.) Proven experience designing and scaling reliability systems for multi-tenant SaaS platforms Strong debugging and systems thinking across distributed microservices and legacy systems Demonstrated ability to lead initiatives that improve incident detection, response, and system resilience Hands-on engineering approach with a track record of building-not just configuring-reliability systems Preferred Qualifications * Experience in B2B SaaS serving enterprise or financial customers * Familiarity with third-party SaaS connector architectures and ingestion patterns * Experience building anomaly detection or intelligent alerting systems * Experience designing customer-facing status pages and incident communication frameworks ## Description As a Sr. Staff SRE at Obsidian, you will define and drive the company-wide reliability vision for a complex, multi-tenant SaaS platform serving enterprise and financial customers. You will operate as a strategic partner to DevOps and Platform Engineering leadership, shaping a unified reliability strategy that scales across the organization. Your core mandate: ensure Obsidian detects, diagnoses, and communicates system issues before customers are impacted-consistently and predictably. This is a hands-on technical role that involves architecting and leading the implementation of systems that handle real-world complexity, including upstream SaaS dependencies, sparse and noisy signals, and mission-critical enterprise workloads., * Reliability Strategy & Architecture - Define and lead long-term reliability strategy across services. Establish end-to-end system visibility frameworks and guide architecture for observability, detection, and resilience. * Cross-Org Leadership - Partner across teams to embed reliability, standardize SLI/SLOs, and serve as a technical escalation expert. * Detection & Observability - Build intelligent detection systems (anomaly detection, connector health models) and enable self-service observability. * Incident Management - Define and evolve a tiered incident communication strategy, improve response practices, and lead postmortems to strengthen reliability and customer trust. * Execution - Contribute hands-on to system design, monitoring, and debugging across distributed systems and data pipelines. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Keycloak case study: Making users happy with service level indicators and observability](https://www.wearedevelopers.com/videos/1599-keycloak-case-study-making-users-happy-with-service-level-indicators-and-observability) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)