> Markdown version of [/jobs/ext/626653-staff-sre](https://www.wearedevelopers.com/jobs/ext/626653-staff-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff SRE - **Company:** Pura, LLC - **Location:** Pleasant Grove, UT, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Architectural Patterns, Cloud Computing, Configuration Management, Information Engineering, Disaster Recovery, Distributed Systems, Fault Tolerance, Identity and Access Management, Python (Programming Language), Node.Js, Performance Tuning, Reliability Engineering, Multi-Cloud, Reliability of Systems, Backend, Kubernetes, Infrastructure Automation Frameworks, Terraform, Programming Languages - **Published:** June 24, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=89c48dec7dea36f2 ## About the Role Do you have experience in Tooling?, * 10+ years of extensive experience as a Site Reliability Engineer or similar role, with a proven track record of architecting solutions for large-scale distributed systems. * Expert-level proficiency in multiple programming languages including Python, Go, or Node.js, with demonstrated experience building complex automation frameworks and infrastructure tools. * Comprehensive mastery of cloud technologies, particularly AWS and GCP, including experience architecting multi-region, highly available systems. * Deep expertise in Kubernetes administration and architecture, including experience operating large-scale clusters, implementing custom controllers, and optimizing cluster performance. * Extensive experience with advanced observability platforms and practices, including implementing custom monitoring solutions and developing sophisticated alerting strategies. * Proven track record of designing and implementing complex IAM architectures for enterprise-scale organizations. * Distinguished expertise in Infrastructure as Code, particularly with Terraform, including experience developing custom providers and managing multi-cloud deployments. * Exceptional problem-solving abilities with demonstrated experience resolving critical production issues in complex, high-stakes environments. ## Description * Architect, design, and implement enterprise-scale infrastructure solutions supporting Web, Mobile, Backend, and Data engineering teams, while providing technical leadership across cross-functional groups. * Define and drive adoption of reliability standards, architectural patterns, and engineering best practices across the organization, working closely with engineering and security leadership. * Lead performance optimization initiatives, implementing sophisticated monitoring strategies and leveraging advanced analytics to ensure exceptional system reliability and performance at scale. * Design and implement comprehensive automation frameworks for infrastructure provisioning, configuration management, and deployment processes, focusing on efficiency and scalability. * Serve as the technical authority for incident management, establishing robust incident response frameworks, leading cross-functional response efforts, and driving systematic improvements through detailed post-incident analysis. * Architect and implement enterprise-wide incident response strategies, including sophisticated playbooks and multi-tier escalation procedures aligned with business continuity requirements. * Partner with engineering leadership to drive reliability improvements through advanced automated testing frameworks, fault-tolerant architectures, and comprehensive disaster recovery strategies. * Provide technical mentorship and leadership to the broader engineering organization while contributing to the strategic direction of the SRE practice. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Terraform for Developers](https://www.wearedevelopers.com/videos/3-terraform-for-developers) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers)