Principal Site Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+6 more
Job description
This is not a Principal role where you are removed from the technology. The team is looking for someone who can remain highly hands-on while also serving as a technical resource and mentor to the rest of the group. Kubernetes is at the center of the environment, with workloads running across EKS and AKS, and future initiatives may involve deeper cluster architecture and design. Today, the focus is heavily around scaling clusters, Kubernetes networking, improving existing architecture, automation, and reliability. The manager strongly values curiosity, attitude, and a willingness to learn over simply checking every technology box, making this a great opportunity for someone who wants continued technical growth while having real influence over the direction of the platform.
Requirements
5+ years of experience within Site Reliability Engineering, DevOps, Platform Engineering, Cloud Engineering, or a similar infrastructure-focused role
Strong hands-on Kubernetes experience
Experience supporting and scaling production Kubernetes clusters
Strong understanding of Kubernetes networking and architecture
Experience with AWS and/or Azure cloud environments
Infrastructure as Code experience, preferably Terraform
Experience supporting highly available, production cloud environments
Strong troubleshooting and incident response experience
Experience with Linux-based environments
Experience working across both SRE and DevOps responsibilities
Ability to mentor engineers and provide technical guidance at a Principal level
Strong communication skills and ability to work across engineering teams
Curiosity and willingness to learn new technologies and cloud platforms
Desired Skills & Experience
Experience with both AWS EKS and Azure AKS
Kubernetes cluster architecture or cluster design experience
Terraform experience in large-scale environments
Python, Go, or another programming/scripting language
Experience building internal tooling or infrastructure automation
Experience with observability, monitoring, SLOs/SLIs, and production reliability
Experience improving Kubernetes scalability, networking, and performance
Exposure to Google Cloud Platform
Experience using AI-assisted engineering tools in a thoughtful and cost-conscious way
Previous technical mentorship or leadership experience
About the company
Our client is a global consumer technology company based in Tempe, Arizona that is looking to add a Principal Site Reliability Engineer to its SRE organization. This is a full-time, hybrid opportunity supporting highly available platforms across AWS and Azure utilizing Kubernetes, Terraform, Linux, cloud-native infrastructure, and modern automation technologies. The team is responsible for both traditional SRE/reliability work and hands-on DevOps/platform engineering, making this a strong opportunity for someone who enjoys owning production systems end to end.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Loading talks and stories from around this role…