Apple Services Engineering (ASE) Compute - Software Engineering Manager

Apple Inc.
Cupertino, CA, United States
2 days ago
Apply on www.techcareers.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Cloud Computing Data Centers Distributed Systems Job Scheduling Python (Programming Language) Kernel-Based Virtual Machine OpenStack Reliability Engineering Ansible Prometheus Software Engineering
+7 more
Data Logging Grafana Kubernetes Infrastructure Automation Frameworks Bare Metal Terraform Dynatrace

Job description

Apple Service Engineering (ASE)’s Compute team is seeking an experienced Software Engineering Manager to lead a team of Infrastructure and Site Reliability Engineers responsible for operating and scaling large-scale batch compute infrastructure across Apple’s data centers. You will manage a team that operates core compute controllers, proxy services, job execution agents, and supporting infrastructure across multiple geographies - ensuring platform availability, reliability, and performance at Apple scale.

You will drive strategic initiatives spanning multi-datacenter capacity planning, incident management, release engineering, observability, and infrastructure modernization. This role requires a leader who can balance operational excellence with engineering innovation, establishing SLOs, driving production readiness, and building the automation and tooling that enable a growing platform to scale efficiently. You will champion the use of AI to accelerate incident triage, improve operational workflows, drive capacity efficiency, and enhance team productivity across all domains.

Requirements

  • 5+ years of experience managing infrastructure, SRE, or platform engineering teams operating large-scale distributed systems
  • Proven track record of building and leading on-call organizations with structured incident management, escalation procedures, and post-incident review processes
  • Strong technical background in cloud infrastructure, compute orchestration, and bare metal provisioning at scale
  • Experience with Kubernetes, OpenStack, KVM/hypervisor technologies, and Infrastructure as Code tools (Chef, Ansible, Terraform, or Salt)
  • Deep understanding of SRE principles including SLOs, error budgets, capacity planning, and release engineering
  • Excellent verbal and written communication skills with the ability to influence across teams and levels
  • Demonstrated ability to recruit, develop, and retain high-performing engineering talent

Preferred Qualifications

  • Hands-on experience leveraging AI and machine learning to improve operational efficiency, incident management, or infrastructure automation
  • Experience managing or scaling batch compute, job scheduling, or HPC platforms
  • Proficiency in Go or Python with a strong automation-first mindset
  • Familiarity with observability stacks (Prometheus, Grafana, distributed tracing) and centralized logging at scale
  • Experience operating large-scale multi-tenant Infrastructure as a Managed Service
  • Experience managing geographically distributed teams and follow-the-sun on-call models
  • Track record of driving capacity efficiency initiatives resulting in measurable cost optimization

About the company

People at Apple don’t just build products - they craft the kind of experience that has revolutionized entire industries. The diverse collection of our people and their ideas inspire innovation in everything we do. Imagine what you could do here! Join Apple, and help us leave the world better than we found it.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.techcareers.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · World Congress 2025

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:09 min

Reevaluating engineering careers at major technology corporations

Chris Heilmann +2 · LIVE

12:08 min

Comparing Keptn orchestration capabilities against alternative software operators

Thomas Schütz · LIVE

Videos

See all

Related articles

See all