Site Reliability Engineer, Apple Data Platform SRE / Apple Services Engineering

Apple Inc.
Cupertino, CA, United States
1 day ago
Apply on www.themuse.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$184,700.0
Working hours
Regular working hours

Tech stack

Airflow Amazon S3 Code Review Computer Programming Data Infrastructure Data Security Linux Disaster Recovery Distributed Data Store Distributed Systems Apache Hadoop Hadoop Distributed File System
+14 more
Apache HBase Python (Programming Language) Reliability Engineering Software Engineering Ceph (Software) Automated Data Processing (ADP) Apache Yarn Apache Spark Generative AI Siri Data Lakes Kubernetes Information Technology Golang

Job description

As a principal contributor and technical lead in our Apple Data Platform (ADP) SRE organization, you will apply SRE principles as you mentor and partner with our engineers and partner teams, ensuring large-scale analytics infrastructure runs reliably and efficiently. This role focuses on driving reliability standards, architectural consistency, and engineering excellence across peer SRE teams and partner engineering organizations - spanning Hadoop, HBase, Spark, Data Lakes, and Airflow ecosystems - through technical leadership, cross-functional alignment, and the development of platform-wide tooling, observability, and operational practices that raise the reliability bar for all of ADP. This role includes production on-call responsibilities., Apple Service Engineering (ASE) teams build and scale the platforms and infrastructure behind many of Apple’s services - including iCloud, iTunes, Siri, and Maps. We are the foundation on which Apple’s software developers build the products that our customers love. We are looking for a passionate and dedicated Technical Lead to drive SRE standards and engineering excellence across the entire Apple Data Platform organization. The Apple Data Platform (ADP) SRE Technical Lead partners with multiple SRE and engineering teams across the data platform - including teams responsible for Hadoop and HBase infrastructure, Spark, S3-compatible storage, and Airflow-orchestrated pipelines. Rather than owning a single vertical, this role sets the technical direction for how reliability is practiced across ADP: defining SLOs, establishing architectural review processes, developing shared tooling and automation, and ensuring that SRE principles are applied consistently as the platform scales. You will be a force multiplier - making every team around you more effective.

Responsibilities:

Serve as the SRE Technical Lead across ADP, partnering with vertical SRE teams and software engineering organizations to ensure reliability standards are consistently applied across the full data platform

Define and drive adoption of SLO frameworks, error budget policies, and incident management practices across ADP services

Provide architectural review and reliability guidance for new services and major platform changes, identifying risks and influencing design before they reach production

Lead the development of shared observability, automation, and infrastructure-as-code tooling that benefits multiple ADP teams simultaneously

Identify and eliminate systemic sources of toil and instability across the platform; advocate for and deliver platform-wide reliability improvements

Mentor and grow SRE engineers across teams, establishing a culture of engineering excellence and continuous improvement

Represent ADP SRE in cross-organizational forums, communicating technical strategy and reliability posture to ASE and Apple leadership

Programming in Python and Golang, supported by Generative AI tooling, to accelerate development of mission-critical shared automation and tools

Production on-call and incident management responsibilities, including leading response for high-severity cross-platform incidents

Requirements

15+ years of experience in SRE or related work managing infrastructure at scale

Experience with Ceph object storage operations

Kubernetes cluster operations experience, particularly running stateful data workloads

Experience with scale testing, disaster recovery, and capacity planning across distributed data systems

Experience driving multi-year platform migrations or large-scale architectural transitions

Ability to define the technical roadmap for a data platform organization and drive cross-functional alignment on architectural standards and best practices

Background in data security, access control, or compliance-sensitive data environments

Minimum Qualifications

BS/MS in Computer Science or equivalent

12+ years of experience in Site Reliability Engineering, managing infrastructure and services at scale

5+ years of experience in technical leadership roles, with demonstrated ability to lead horizontally across teams without direct authority

Broad expertise across the data platform stack: Hadoop (HDFS, YARN), HBase, Apache Spark, Data Lake architectures, S3-compatible storage solutions, and Apache Airflow

History of defining and driving SLO/error budget frameworks and reliability practices across multiple teams or services

Demonstrable programming skills to develop shared tooling, lead code reviews, and set engineering standards, Strong written and verbal communication skills - able to present technical strategy to both engineers and leadership

Advanced knowledge of Linux, networking, and distributed systems fundamentals

Benefits & conditions

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $184,700 and $324,800, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses - including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits

About the company

At Apple, we believe that innovation flourishes in an environment where ideas are challenged, collaboration is encouraged, and technology is pushed to its limits. This environment is only possible when diverse minds come together, bringing unique perspectives and experiences. Our people and their ideas inspire innovation in everything we do. Imagine what you could accomplish here! Join Apple and help us make the world a better place.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.themuse.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:16 min

Evaluating the enduring financial and technological legacy of Apple

Marco Landi · World Congress 2024

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all