> Markdown version of [/jobs/ext/1057688-staff-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1057688-staff-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Site Reliability Engineer - **Company:** Ironclad, Inc. - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $220,000.0 - $235,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Cloud Storage, Computer Programming, Cursor (Graphical User Interface Elements), DevOps, Distributed Systems, Octopus Deploy, Reliability Engineering, TypeScript, AI Infrastructure, Circleci, Pulumi, Google Cloud, Cloud Platform System, Kubernetes, Terraform - **Published:** June 9, 2026 - **Apply:** https://www.dice.com/job-detail/53256f52-ff46-43bc-a09f-46d0e9f111f8 ## About the Role * 8+ years of DevOps / SRE experience * 5+ years of coding experience * Expert knowledge of Kubernetes and Google Cloud Platform (or similar provider) * Ability to build resilient infrastructure * Modern GitOps - Experience with tools like Terraform/Pulumi, CircleCI, ArgoCD * Experience with modern AI enabled tools such as Claude Code, Cursor, Zed * Troubleshooting and analytical skills, can PR review human and AI generated code * Strong technical aptitude and exceptional communication skills (written and verbal) * Desire for helping customers, and the ability to dive deep and learn a new product. * Experience and desire to work cross-functionally * Team and goal-oriented. * High output; low ego Bonus Points if you have: * Experience with multi-region support * Expert Database Management Experience * Experience managing AI Infrastructure * Typescript Experience ## Description This is a hybrid role. Office attendance is required at least twice a week on Tuesdays and Thursdays for collaboration and connection. There may be additional in-office days for team or company events. We are seeking a strategic, high-output Staff/Senior Staff SRE to define the future of our cloud platform and champion engineering excellence across Ironclad. In this role, you will pair deep cloud-native mastery with a passion for mentoring, driving architectural resilience, and solving complex distributed systems at scale. Roles & Responsibilities: * Provide technical leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud Platform * Define and champion SRE best practices, setting the standard for engineering excellence across the entire organization * Solve the whole problem. Architecture for resiliency, identify risks, and make it happen. * A proven track record of designing and driving an 'automate-everything' culture (build, test, deploy, monitor) * Preference for collaboration, open communication and reaching across functional borders * Thorough understanding of backup/recovery systems, cloud storage architecture, and distributed systems * Be on an on-call rotation to respond to incidents that impact Ironclad's availability, and provide support with internal or customer-facing incidents * Translate the near, mid, and long-term strategic needs of the business into a scalable, resilient platform roadmap * Drive critical architectural decisions with a relentless focus on security, scalability, and high performance * Be a mentor, multiply our team's output with leadership and guidance ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Securing Your Web Application Pipeline From Intruders](https://www.wearedevelopers.com/videos/53-securing-your-web-application-pipeline-from-intruders) - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) - [Serverless: Past, Present and Future](https://www.wearedevelopers.com/videos/34-serverless-past-present-and-future) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Dev Digest 162: AI careers, MCP, AWS best practices & floppy sweaters](https://www.wearedevelopers.com/magazine/571-dev-digest-162-ai-careers-mcp-aws-best-practices-floppy-sweaters)