> Markdown version of [/jobs/ext/1180097-principal-site-reliability-engineer-austin-texas](https://www.wearedevelopers.com/jobs/ext/1180097-principal-site-reliability-engineer-austin-texas). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Site Reliability Engineer - Austin, Texas - **Company:** ZOWTA, LLC - **Location:** Austin, TX, United States - **Experience:** Expert - **Salary:** $93,309.0 - $143,208.0 - **Contract:** Permanent contract - **Skills:** Agile Methodology, Artificial Intelligence, Amazon Web Services, Software as a Service, Cloud Computing, Cloud Computing Security, Cloud Engineering, DevOps, Linux System Administration, Reliability Engineering, Software Engineering, Data Logging, Cloud Platform System, System Availability, Gitlab, Containerization, Kubernetes, Production Code, Build Tools, Terraform, Dynatrace - **Published:** July 4, 2026 - **Apply:** https://www.careerjet.com/jobad/us7ea3b6a43153b53f1afcb65eb7fd4f05 ## About the Role * 10+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, Cloud Infrastructure, or Software Engineering. * Proven experience designing and operating large-scale, highly available cloud infrastructure in AWS. * Strong software engineering background with the ability to write production-quality code and automation. * Expert-level experience with Infrastructure as Code, preferably Terraform. * Deep experience designing and maintaining modern CI/CD pipelines using GitLab or similar platforms. * Strong knowledge of Kubernetes, containerized workloads, and cloud-native architectures. * Extensive experience with observability platforms, distributed tracing, logging, monitoring, and incident response. * Experience defining and implementing SLOs, SLIs, and reliability engineering best practices. * Strong understanding of networking, security, Linux systems administration, and cloud architecture. * Experience supporting high-traffic SaaS applications and mission-critical production environments. * Excellent problem-solving skills with the ability to simplify complex technical challenges. * Demonstrated ability to influence technical direction without direct authority while mentoring engineers across multiple teams. * Experience working in Agile development environments and partnering closely with cross-functional engineering teams. ## Description ShipperHQ is looking for a Principal Site Reliability Engineer to lead the evolution of our cloud platform, reliability strategy, and infrastructure architecture. This is a highly technical, hands-on leadership role responsible for designing scalable, resilient systems while establishing engineering best practices that enable our teams to move quickly and confidently. As a Principal SRE, you'll own the strategic direction of our cloud infrastructure, deployment architecture, observability, and platform reliability. You'll partner closely with Engineering, Product, Security, and QA to build systems that are secure, automated, highly available, and built to scale. Success in this role comes from balancing strategic thinking with execution and leading through influence, solving complex technical challenges, and continuously improving the developer experience. This role is ideal for someone who enjoys building platforms rather than simply maintaining infrastructure and thrives in a fast-paced, AI-first engineering culture. * Own the technical vision and roadmap for ShipperHQ's cloud infrastructure, reliability, and platform engineering initiatives. * Design, build, and maintain highly available, scalable, and secure cloud infrastructure in AWS. * Architect and evolve Infrastructure as Code (Terraform) standards across all environments. * Design and optimize CI/CD pipelines that enable fast, reliable, and repeatable software delivery. * Define and implement reliability standards, SLOs, SLIs, error budgets, and incident management best practices. * Lead the design and implementation of observability, monitoring, logging, and alerting across the platform. * Build self-service platform capabilities and automation that empower engineering teams and reduce operational overhead. * Drive infrastructure modernization initiatives, including containerization, orchestration, and platform scalability. * Partner with Security to implement cloud security best practices, compliance controls, and governance. * Collaborate with Engineering teams to improve application reliability, performance, and operational excellence. * Lead technical decision-making for infrastructure architecture and serve as a trusted advisor across engineering. * Mentor engineers and promote best practices in cloud architecture, automation, reliability, and operational excellence. * Evaluate and introduce new technologies that improve scalability, reliability, developer productivity, and operational efficiency. * Participate in incident response, root cause analysis, and continuous improvement efforts for production systems. ## Related Videos - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [The Power of Purpose: Unlocking Potential and Innovation](https://www.wearedevelopers.com/videos/1110-the-power-of-purpose-unlocking-potential-and-innovation) - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers)