> Markdown version of [/jobs/ext/2706827-principal-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2706827-principal-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Site Reliability Engineer - **Company:** Tandem Diabetes Care - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $165,000.0 - $185,000.0 - **Contract:** Temporary contract - **Skills:** Agile Methodology, Amazon Web Services, Microsoft Azure, Bash Shell, Cloud Computing, Cloud Computing Security, Code Review, Continuous Integration, DevOps, Disaster Recovery, Failover, Github, Identity and Access Management, Python (Programming Language), Load Testing, Network Segmentation, Octopus Deploy, Reliability Engineering, Cloud Services, Prometheus, Software Engineering, Software Vulnerability Management, Backup and Restore, Datadog, Data Logging, Scripting, Cloud Platform System, Grafana, Mttr, Containerization, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Patch Management, Cloud Optimization, Cloudwatch, Terraform, New Relic (SaaS), Docker, Pagerduty, Golang, Programming Languages - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/principal-site-reliability-engineer-temp-to-hire-tandem-diabetes-care-interna-9796552 ## About the Role \* Demonstrated experience leading production support and incident management for production systems, including incident command during high-severity events. \* Strong grounding in SRE principles: SLIs/SLOs, blameless postmortems, toil reduction, and treating reliability as an engineering discipline rather than a purely operational function. \* Demonstrated experience owning on-call strategy, including rotation design, alert tuning, and escalation. \* Expertise with Terraform or comparable IaC at scale: module design, state management, and policy-as-code guardrails. \* Hands-on experience building CI/CD pipelines with reliability guardrails using tools such as GitHub Actions, Octopus Deploy, or Azure DevOps. \* Deep experience with at least one major cloud platform (AWS, Azure, or GCP; [ ## Description The Principal Site Reliability Engineer (SRE) is responsible for the reliability, availability, and performance of the company's production systems. This role leads day-to-day production support and incident response and progressively replaces reactive firefighting with engineered SRE practice: SLOs, observability, on-call design, runbooks, and automation. It also advances infrastructure automation (Terraform/IaC), CI/CD reliability, disaster recovery, and security and compliance readiness in partnership with cross-functional teams. The role works with a distributed team that includes offshore consulting partners and is expected to raise their capability and independence; success is measured as much by what the team can do without the Principal SRE as by what the Principal SRE delivers personally. Principal Site Reliability Engineer's at Tandem are also responsible for: Production Support & Incident Management * Leads day-to-day production support: intake, triage, prioritization, escalation, queue health, and change execution. * Establishes consistent support practices across a distributed team, including shift handoffs, ticket quality standards, and clear ownership of open issues. * Leads incident management end-to-end: incident command, stakeholder communication, and blameless postmortems with corrective actions tracked to closure. * Participates in and coordinates response activities for production security incidents, partnering with Security teams to contain threats, restore services, validate controls, and implement corrective actions. * Owns on-call strategy: rotation design, escalation paths, alert tuning, and tooling (e.g., PagerDuty, New Relic), with explicit goals of reducing alert fatigue and building coverage that works across time zones. Participates as a senior escalation tier for high-severity incidents. * Builds runbooks that standardize response to common failure modes and enable first-line resolution by engineers who did not build the system. Reliability Engineering & Observability * Defines and owns SLIs and SLOs for critical services, using them to guide monitoring, alerting, and reliability priorities. * Drives systemic reduction of MTTD and MTTR through better instrumentation, alerting, diagnostics, and automation. * Converts recurring support burden into permanent fixes, automation, or documentation rather than absorbing it as ongoing manual work. * Establishes and maintains technology currency and lifecycle management practices for production platforms, ensuring cloud services, Kubernetes clusters, operating systems, runtimes, and infrastructure components remain supported, secure, and aligned with organizational standards. Proactively identifies and mitigates End-of-Life (EOL), End-of-Support (EOS), and technology obsolescence risks. * Owns business continuity and disaster recovery readiness for production platforms, including backup and recovery strategies, recovery testing, failover capabilities, recovery runbooks, and adherence to defined RTO/RPO objectives. Infrastructure Automation & CI/CD * Leads infrastructure automation with Terraform, ensuring infrastructure is version-controlled, modular, reusable, and auditable, working within and improving existing patterns where they are sound. * Eliminates toil through automation, reducing manual, repetitive operational work across the team. * Adds reliability guardrails to CI/CD pipelines (automated rollback, change-risk checks, progressive delivery) with the teams that own them. Operational Governance and Compliance * Maintains production systems in accordance with applicable regulatory and compliance requirements, ensuring audit readiness and the ongoing effectiveness of access management, change management, logging, vulnerability remediation, patch management, and other operational controls. * Maintains documentation and evidence supporting business continuity and disaster recovery controls, including recovery testing results, remediation plans, and audit artifacts. * Partners with Security, Quality, and Compliance teams to maintain regulatory compliance, support internal and external audits, and drive timely resolution of audit findings and remediation activities. Leadership, Mentorship & Collaboration * Actively grows the capability of SRE and DevOps engineers, including consulting-partner engineers, through pairing, design and code review, and incident debriefs. * Creates conditions in which junior and contract engineers raise concerns and disagree openly, recognizing that silent agreement is a production risk. * Communicates in a documentation-first manner suited to a distributed, multi-time-zone team, so context is durable rather than held by one person. * Partners with software engineering, QA, and architecture to embed reliability and operability into the development lifecycle rather than only after incidents. Contributing Responsibilities (partners with other owners): * Informs capacity planning and scaling strategy with the Test team (who own load testing and modeling) and introduces proactive resilience testing such as game days once baseline support and observability are stable. * Partners with application, architecture, and business stakeholders to align business continuity and disaster recovery capabilities with application requirements and recovery objectives. * Supports cloud cost optimization: rightsizing, reserved capacity, and observability spend governance. * Ensures work complies with company policies, including Privacy/HIPAA and other regulatory, legal, and safety requirements. Other duties as assigned. WHEN & WHERE YOU'LL WORK: Remote: This position is fully remote and open to candidates within the United States. Equipment for the role will be provided and training will occur virtually. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025)