> Markdown version of [/jobs/ext/2081360-apple-services-engineering-ase-compute-software-engineering-manager](https://www.wearedevelopers.com/jobs/ext/2081360-apple-services-engineering-ase-compute-software-engineering-manager). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Apple Services Engineering (ASE) Compute - Software Engineering Manager - **Company:** Apple Inc. - **Location:** Cupertino, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Cloud Computing, Data Centers, Distributed Systems, Job Scheduling, Python (Programming Language), Kernel-Based Virtual Machine, OpenStack, Reliability Engineering, Ansible, Prometheus, Software Engineering, Data Logging, Grafana, Kubernetes, Infrastructure Automation Frameworks, Bare Metal, Terraform, Dynatrace - **Published:** August 16, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3355987610&tx=UT545THZ&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * 5+ years of experience managing infrastructure, SRE, or platform engineering teams operating large-scale distributed systems * Proven track record of building and leading on-call organizations with structured incident management, escalation procedures, and post-incident review processes * Strong technical background in cloud infrastructure, compute orchestration, and bare metal provisioning at scale * Experience with Kubernetes, OpenStack, KVM/hypervisor technologies, and Infrastructure as Code tools (Chef, Ansible, Terraform, or Salt) * Deep understanding of SRE principles including SLOs, error budgets, capacity planning, and release engineering * Excellent verbal and written communication skills with the ability to influence across teams and levels * Demonstrated ability to recruit, develop, and retain high-performing engineering talent Preferred Qualifications * Hands-on experience leveraging AI and machine learning to improve operational efficiency, incident management, or infrastructure automation * Experience managing or scaling batch compute, job scheduling, or HPC platforms * Proficiency in Go or Python with a strong automation-first mindset * Familiarity with observability stacks (Prometheus, Grafana, distributed tracing) and centralized logging at scale * Experience operating large-scale multi-tenant Infrastructure as a Managed Service * Experience managing geographically distributed teams and follow-the-sun on-call models * Track record of driving capacity efficiency initiatives resulting in measurable cost optimization ## Description Apple Service Engineering (ASE)'s Compute team is seeking an experienced Software Engineering Manager to lead a team of Infrastructure and Site Reliability Engineers responsible for operating and scaling large-scale batch compute infrastructure across Apple's data centers. You will manage a team that operates core compute controllers, proxy services, job execution agents, and supporting infrastructure across multiple geographies - ensuring platform availability, reliability, and performance at Apple scale. You will drive strategic initiatives spanning multi-datacenter capacity planning, incident management, release engineering, observability, and infrastructure modernization. This role requires a leader who can balance operational excellence with engineering innovation, establishing SLOs, driving production readiness, and building the automation and tooling that enable a growing platform to scale efficiently. You will champion the use of AI to accelerate incident triage, improve operational workflows, drive capacity efficiency, and enhance team productivity across all domains. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [The Power of Purpose: Unlocking Potential and Innovation](https://www.wearedevelopers.com/videos/1110-the-power-of-purpose-unlocking-potential-and-innovation) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Eclipse Che for Infrastructure Automation](https://www.wearedevelopers.com/videos/1611-eclipse-che-for-infrastructure-automation) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)