SRE Software Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+24 more
Job description
The ASE Compute team is looking for a Site Reliability Engineer to deploy and manage a large Kubernetes platform that Apple’s services run on, partnering with engineering teams across the company to solve complex problems using both open-source and in-house tooling. You will contribute to the development of our controllers and namespace management infrastructure, working alongside senior engineers to strengthen the reliability of our Kubernetes services. You will learn to write well-tested code, participate in design reviews, and gradually take ownership of features. You’ll have the opportunity to engage with the upstream community, gain hands-on experience with production-scale systems, and build the technical foundation to support service teams across Apple. The role also offers room to build AI-assisted tooling that accelerates triage, operational workflows, and infrastructure automation for the whole team., Deploy, configure, and maintain large-scale, multi-tenant Kubernetes environments
Write and maintain operational tooling to improve reliability and reduce manual intervention
Implement and maintain reliability standards for the platform: SLOs, error budgets, alerting philosophy, upgrade and rollout strategy, and the run-books that follow from them.
Contribute to CI/CD pipelines, revision control workflows, and configuration management practices
Take on-call, troubleshoot production issues, and follow up on post-incident action items
Help enforce security best practices, OS hardening, and compliance standards across the fleet
Requirements
Site Reliability Engineering, DevOps, or Infrastructure focused experience
Experience with third-party cloud platforms (AWS, GCP, or Azure)
Experience with containerization and orchestration technologies such as Docker or Kubernetes
Familiarity with bare-metal provisioning and lifecycle management at datacenter scale
Understanding of cloud-native observability (Prometheus, Thanos, Splunk, or similar)
Familiarity with CI/CD pipelines and DevOps practices
Knowledge of OS security hardening, encryption, and regulatory compliance frameworks
Minimum Qualifications
Hands-on experience in Linux systems administration and containerization with enterprise distributions such as RHEL, Oracle Linux, or CentOS
Proficiency in Python or Go for scripting and tooling
Solid understanding of Linux fundamentals: file systems, process management, user and group administration, and package managementWorking knowledge of networking concepts including TCP/IP, DNS, DHCP, and basic firewall configuration
Experience with version control systems such as Git and configuration management (Puppet, Ansible, or equivalent)
Strong written and verbal communication skills
About the company
People at Apple don’t just build products, they craft the kind of experience that has revolutionized entire industries. The diverse collection of our people and their ideas inspire innovation in everything we do. Imagine what you could do here! Join Apple, and help us leave the world better than we found it.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Is Software Engineering Over-Saturated?
Dev Digest 120 - Apple and peers
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Now is the time for industrialized software development