Site Reliability Engineering (SRE) Manager

Apple Inc.
Cupertino, CA, United States
2 days ago
Apply on www.themuse.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Compensation
$267,800.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Apple Maps Systems Engineering Cloud Computing Linux Distributed Systems Reliability Engineering Large Language Models Kubernetes

Job description

We are looking for a senior SRE leader to set the strategic direction for our Maps serving infrastructure. This is not just a ā€œkeep the lights onā€ role - this is a leadership position that defines where our infrastructure goes next, how our SRE practice evolves, and how we build the teams and partnerships to get there. You will work closely with engineering, product, and operations partners across Apple to shape the roadmap for our serving platform. You will lead an organization of SREs and hold responsibility for the reliability, scalability, and operational excellence of some of Apple’s most visible services. We believe AI will fundamentally reshape how SRE is practiced - from incident detection and resolution to capacity planning and toil elimination - and we’re looking for a leader who shares that conviction and can drive that transformation across the organization.

Responsibilities:

Operate with the fundamental principle that reliability is feature number one

Define and drive the strategic roadmap for Apple Maps serving infrastructure in partnership with engineering and product teams

Lead and grow an SRE organization, setting the bar for technical excellence, operational rigor, and engineering culture

Represent the SRE perspective in cross-functional planning - translating reliability requirements into architecture decisions and investment priorities

Champion the ā€œEngineeringā€ in Site Reliability Engineering - driving automation, platform improvements, capacity strategy, and systems design beyond reactive incident response

Define and execute a clear AI strategy for the SRE organization - identifying high-impact opportunities where AI/ML-powered tooling can reduce toil, accelerate root cause analysis, and improve reliability outcomes

Drive adoption of AI-assisted tooling (copilots, intelligent runbooks, LLM-based diagnostics, anomaly detection) into day-to-day SRE workflows

Build a culture where engineers actively experiment with AI tools and modern approaches to solve operational problems, Build strong partnerships across Apple, negotiating priorities and aligning on shared goals with an Apple-first mindset

Communicate effectively at the executive level - presenting strategy, trade-offs, and progress to senior leadership

Mentor and develop leaders within your organization, creating a culture where people do their best work

Celebrate wins, recognize contributions, and make tough calls when needed

Requirements

Experience with cloud infrastructure (AWS, GCP) and Kubernetes at scale

Background in capacity planning, performance engineering, or infrastructure architecture

Track record of driving cultural and process transformation within SRE organizations

Experience building or deploying AI-powered operational tooling (AIOps, intelligent alerting, automated diagnostics)

Hands-on experience with LLM-based developer/SRE productivity tools

Track record of driving AI adoption within engineering teams

Experience operating services at Apple-scale user volumes

Minimum Qualifications

10-15+ years of experience in SRE or adjacent disciplines (systems engineering, infrastructure engineering, production engineering), with at least 10 years in senior management roles

Demonstrated experience leading SRE organizations supporting large-scale, user-facing distributed services

Strong technical proficiency in Linux fundamentals, distributed systems concepts, networking, and infrastructure at scale

Demonstrated experience applying AI/ML tooling or LLM-based solutions to improve SRE or infrastructure operations

Ability to read and understand code produced by LLMs and evaluate its suitability for production use

Has defined or is actively executing an AI strategy for a large SRE organization

Proven ability to communicate at the executive level and negotiate across organizational boundaries

Benefits & conditions

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $267,800 and $401,700, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses - including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.themuse.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter Ā· World Congress 2022

5:34 min

Industry examples of disastrous technology and game launches

Brenda Romero Brenda Romero Ā· World Congress 2023

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard Ā· World Congress 2025

3:09 min

Reevaluating engineering careers at major technology corporations

Chris Heilmann +2 Ā· LIVE

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn Ā· World Congress 2022

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou Ā· Coffee With Developers

Videos

See all

Related articles

See all