Site Reliability Engineer, Apple Data Platform Compute SRE

Apple Inc.
United States
9 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

ITunes Code Review Computer Programming Data Infrastructure Linux Disaster Recovery Distributed Data Store Apache Hadoop ICloud Reliability Engineering Software Engineering Automated Data Processing (ADP)
+4 more
Siri Kubernetes Information Technology Bare Metal

Job description

As a principal contributor in our Apple Data Platform SRE organization you will apply SRE principles as you mentor and partner with our engineers and partner teams, ensuring petabyte-scale analytics infrastructure runs reliably and efficiently. This role focuses on managing bare-metal and cloud based infrastructure, levering and extending our infrastructure-as-code based tooling, analyzing and optimizing performance, helping to plan and execute long term fleet management logistics, capacity planning, and ultimately maintaining operational excellence across distributed data platforms that power analytics across Apple. This role includes production on-call responsibilities., Apple Service Engineering (ASE) teams build and scale the platforms and infrastructure behind many of Apple’s services (such as iCloud, iTunes, Siri, and Maps). We are the foundation on which Apple’s software developers build the products that our customers love. We are looking for a passionate and dedicated Senior Site Reliability Engineer to provide technical leadership on our team to help ensure our customers have the highest quality Apple Services experience. The Apple Data Platform (ADP) Compute SRE team is responsible for the core infrastructure, including our legacy bare-metal platforms and modern cloud based infrastructure stack. We partner with both peer SRE teams and several of our world-class software and product engineering teams to support infrastructure reliability, multi-year parallel migrations for Apple properties, as well as the automation, tooling, incident, and process management necessary to ensure smooth 24x7 operations for ADP customers.

Requirements

12+ years of experience in Site Reliability Engineering, managing infrastructure and services at scale

5+ years of experience in management or technical leadership roles

History of end-to-end project management and delivery

Demonstrable programming skills to both develop software/tools and lead code reviews

Experience managing Hadoop and Kubernetes infrastructure and related services, or equivalent experience

Advanced knowledge of Linux, Networking, and Containers

Preferred Qualifications

15+ YoE in SRE or related work managing infrastructure at scale

Experience with scale testing, disaster recovery, and capacity planning

Ability to define the technical roadmap for infrastructure and drive cross-functional alignment on architectural standards and best practices

About the company

At Apple, we believe that innovation flourishes in an environment where ideas are challenged, collaboration is encouraged and technology is pushed to its limits. This environment is only possible when diverse minds come together, bringing unique perspectives and experiences. Our people and their ideas inspire innovation in everything we do. Imagine what you could accomplish here! Join Apple and help us make the world a better place.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:06 min

Developer experience and project variety at scale

Alexandra Petri · WWC 2023

1:16 min

Evaluating the enduring financial and technological legacy of Apple

Marco Landi · WWC 2024

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

2:23 min

Historical breakthroughs in natural language processing models

Mary Grygleski Mary Grygleski · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

5:00 min

Exploring the specific workplace responsibilities of staff software engineers

Jan Giacomelli · LIVE

Videos

See all

Related articles

See all