> Markdown version of [/jobs/ext/2716688-staff-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2716688-staff-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Site Reliability Engineer - **Company:** Pivotal Health, Inc. - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Audit Trail, Cloud Computing, Cloud Engineering, Software Debugging, Disaster Recovery, Distributed Systems, Reliability Engineering, Site Reliability Engineering Practices, Delivery Pipeline, Data Layers, Infrastructure Automation Frameworks, Build Tools, Machine Learning Operations - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/staff-site-reliability-engineer-pivotal-health-9610674 ## About the Role * 8+ years of experience in site reliability engineering, infrastructure engineering, platform engineering, or operating large-scale production systems * Deep knowledge of cloud infrastructure, distributed systems, networking, containers, orchestration, infrastructure as code, and modern deployment practices * Experience designing, building, and operating highly available systems in a fast-growing production environment * Strong understanding of observability, service-level objectives, capacity planning, incident management, disaster recovery, and performance engineering; Hands-on and technically credible, with the ability to debug complex issues across application, infrastructure, network, and data layers * Experienced in creating automation and internal tooling that reduces operational toil and improves developer productivity * Comfortable influencing architecture and engineering practices across teams without relying on formal authority * A thoughtful communicator who can translate operational risk and technical tradeoffs for engineering, product, security, and business stakeholders; Pragmatic about balancing reliability, delivery speed, complexity, and cost * Comfortable operating in ambiguity and building from scratch-you're energized by greenfield work, not slowed down by it Extra credit if you have * Experience operating infrastructure that handles healthcare, financial, or other sensitive and regulated data * Familiarity with HIPAA, PHI, SOC 2, data privacy, auditability, and security controls * Experience supporting data-intensive, AI, or machine learning systems in production * Experience with multi-region architecture, high-volume asynchronous workflows, or complex distributed processing systems * A track record of introducing SRE practices at an early-stage or high-growth company, If you're excited by solving complex problems and making a real-world impact, we'd love to hear from you., Candidates must be authorized to work in the United States without current or future employer sponsorship. ## Description We're hiring a Staff Site Reliability Engineer to define and strengthen how reliability, scalability, and operational excellence are built into Pivotal's platform. This is a senior individual contributor role with broad influence across engineering. You'll work hands-on with our infrastructure and production systems while setting the technical direction for reliability across the company. You'll help us evolve from a rapidly growing platform into one that can scale predictably, recover gracefully, and meet the high standards of availability, security, and auditability required in healthcare. You'll partner closely with software, data, AI, security, and product teams to design resilient systems, improve observability, reduce operational risk, and make reliability a shared engineering responsibility. You'll bring the technical depth to solve our most difficult infrastructure challenges and the leadership to establish practices that raise the bar for the entire organization. If you're energized by complex distributed systems, greenfield infrastructure work, and the opportunity to shape reliability at a pivotal stage of company growth, this is the role for you., * Set Pivotal's reliability strategy: Define the technical vision and roadmap for reliability, availability, scalability, and operational readiness. Establish clear service-level objectives and help teams make informed tradeoffs between reliability, velocity, and cost. * Design resilient, scalable infrastructure: Guide and implement improvements to our cloud architecture, deployment systems, networking, compute, storage, and other shared infrastructure. Ensure our systems can scale with increasing product usage, data volume, and workflow complexity. * Build world-class observability: Develop a cohesive approach to metrics, logs, traces, dashboards, and alerting. Give engineers the visibility they need to understand system behavior, identify emerging issues, and resolve production incidents quickly. * Improve incident response and resilience: Establish effective incident management, on-call, postmortem, and disaster recovery practices. Lead the response to complex incidents and ensure lessons result in durable improvements to our systems and processes. * Reduce operational toil through automation: Identify recurring manual work and build systems, tooling, and automation that make operating Pivotal's platform safer and more efficient. Improve deployment workflows, capacity management, infrastructure provisioning, and production diagnostics. * Embed reliability across engineering: Partner with software, data, and AI engineers to improve system design, production readiness, and failure handling. Create reusable patterns, tooling, and standards that allow teams to build reliable services without becoming dependent on a centralized operations function. * Strengthen security and compliance: Work closely with security and compliance stakeholders to protect sensitive healthcare and financial data. Help ensure our infrastructure, access controls, audit trails, and operational practices meet applicable regulatory and customer requirements. * Provide technical leadership: Serve as a trusted technical partner to senior engineers and engineering leaders. Lead architecture reviews, mentor engineers, clarify complex tradeoffs, and raise the standard for infrastructure and operational engineering across the organization. ## Related Videos - [Resilient by Design: Building Robust Architectures in High-Stakes Financial Systems](https://www.wearedevelopers.com/videos/2106-resilient-by-design-building-robust-architectures-in-high-stakes-financial-systems) - [Why Your AI Agent Keeps Hallucinating Your Data: Building Deterministic Context Layers](https://www.wearedevelopers.com/videos/2055-why-your-ai-agent-keeps-hallucinating-your-data-building-deterministic-context-layers) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [How to govern Vibe Coding for the Enterprise](https://www.wearedevelopers.com/videos/100290-how-to-govern-vibe-coding-for-the-enterprise) - [No Keys for the Robot: GitOps as the Control Plane for Autonomous Agents](https://www.wearedevelopers.com/videos/100095-no-keys-for-the-robot-gitops-as-the-control-plane-for-autonomous-agents) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)