Principal Site Reliability Engineer

Motion Recruitment Partners LLC.
Tempe, AZ, United States
15 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Cloud Engineering DevOps Python (Programming Language) Linux System Administration Reliability Engineering Software Tools Scripting Google Cloud
+6 more
Cloud Platform System Kubernetes Infrastructure Automation Frameworks Azure AKS Terraform AWS EKS

Job description

This is not a Principal role where you are removed from the technology. The team is looking for someone who can remain highly hands-on while also serving as a technical resource and mentor to the rest of the group. Kubernetes is at the center of the environment, with workloads running across EKS and AKS, and future initiatives may involve deeper cluster architecture and design. Today, the focus is heavily around scaling clusters, Kubernetes networking, improving existing architecture, automation, and reliability. The manager strongly values curiosity, attitude, and a willingness to learn over simply checking every technology box, making this a great opportunity for someone who wants continued technical growth while having real influence over the direction of the platform.

Requirements

5+ years of experience within Site Reliability Engineering, DevOps, Platform Engineering, Cloud Engineering, or a similar infrastructure-focused role

Strong hands-on Kubernetes experience

Experience supporting and scaling production Kubernetes clusters

Strong understanding of Kubernetes networking and architecture

Experience with AWS and/or Azure cloud environments

Infrastructure as Code experience, preferably Terraform

Experience supporting highly available, production cloud environments

Strong troubleshooting and incident response experience

Experience with Linux-based environments

Experience working across both SRE and DevOps responsibilities

Ability to mentor engineers and provide technical guidance at a Principal level

Strong communication skills and ability to work across engineering teams

Curiosity and willingness to learn new technologies and cloud platforms

Desired Skills & Experience

Experience with both AWS EKS and Azure AKS

Kubernetes cluster architecture or cluster design experience

Terraform experience in large-scale environments

Python, Go, or another programming/scripting language

Experience building internal tooling or infrastructure automation

Experience with observability, monitoring, SLOs/SLIs, and production reliability

Experience improving Kubernetes scalability, networking, and performance

Exposure to Google Cloud Platform

Experience using AI-assisted engineering tools in a thoughtful and cost-conscious way

Previous technical mentorship or leadership experience

About the company

Our client is a global consumer technology company based in Tempe, Arizona that is looking to add a Principal Site Reliability Engineer to its SRE organization. This is a full-time, hybrid opportunity supporting highly available platforms across AWS and Azure utilizing Kubernetes, Terraform, Linux, cloud-native infrastructure, and modern automation technologies. The team is responsible for both traditional SRE/reliability work and hands-on DevOps/platform engineering, making this a strong opportunity for someone who enjoys owning production systems end to end.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Loading talks and stories from around this role…