> Markdown version of [/jobs/ext/2209108-software-development-engineer-compute-platform](https://www.wearedevelopers.com/jobs/ext/2209108-software-development-engineer-compute-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Development Engineer, Compute Platform - **Company:** Apple Inc. - **Location:** Seattle, WA, United States - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Data Analysis, Apple Products, Computing Platforms, C++ (Programming Language), Code Review, Computer Programming, Computer Engineering, Software Debugging, Distributed Systems, Python (Programming Language), Operational Databases, Software Engineering, Kubernetes, Information Technology, Slurm, Golang - **Published:** August 24, 2026 - **Apply:** https://www.seattlejobs.com/job.asp?id=3363835705&tx=JJ10195LFR&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * Strong programming skills in a general-purpose language (Go, Python, C++, or similar), and demonstrated ability to design, debug, and test complex software systems * Working knowledge of distributed systems and operating system fundamentals * Experience instrumenting large-scale infrastructure and analyzing the data it produces - reasoning quantitatively about how CPU, memory, I/O or network consumption behaves and how it scales - and turning that analysis into a decision or a measurable improvement in production * Comfort reasoning quantitatively about production data - distributions and tails, not just averages - and about the risk a wrong estimate creates for a running workload * Track record of owning a project from an ambiguous problem statement through production * Strong communication and organizational skills, including the ability to work directly with customers and to make quantitative results legible and actionable for people who are not data specialists * BS in Computer Science / related fields and 3+ years of experience or MS with 1+ years of experience or PhD, * Experience building or operating a large-scale compute platform with responsibility for its capacity, efficiency, or performance * Experience improving utilization on a production platform - right-sizing requests from historical usage, oversubscription, bin packing, co-locating batch alongside latency-sensitive work, or reclaiming unallocated capacity. Comparable work on Kubernetes VPA, on runtime or demand prediction feeding an HPC scheduler, or on an in-house equivalent is equally relevant * Experience characterizing workloads at fleet scale, for example grouping jobs into behavioral classes or building the telemetry to do so * Experience shipping a model or heuristic that made automated decisions in production, and owning the outcome * Experience with forecasting, regression, or uncertainty estimation applied to operational time series * Experience modelling how a system's resource consumption scales - for example projecting the network, storage or I/O demand created by growing a compute footprint - and using that to inform capacity or sizing decisions * Experience with Kubernetes / Slurm or a comparable batch scheduling system ## Description In this role, you will develop, debug, and maintain data-driven features of a large-scale batch focussed compute platform. You will: * Own features end to end that turn job behavior into decisions the platform acts on - from the initial data analysis through the model or heuristic, the APIs, production serving, and the measurement that proves it worked * Implement your own changes across the stack, from platform services and control plane paths down to node-level resource limits and configuration * Design and run validation in production: shadow mode, staged rollouts, explicit success metrics, and a rollback plan * Have full ownership for the decisions your models make. When a prediction misbehaves, diagnose it, bound the impact, and help fix the system that trusted it * Write and review code, generate and review design documentation * Engage directly with customers and partner teams on their compute needs: understand their workloads, analyze cluster utilization and capacity, and turn that analysis into demand forecasts, node pool configuration and quota decisions that balance customer demand against fleet efficiency * Bring quantitative analysis to open questions across the team - sizing, planning and prioritization decisions where good data changes the answer * Participate in software qualifications and rollouts to production clusters * Participate in an on-call rotation where engineers respond to platform issues for same-day resolution * Work with a wide range of software and hardware engineering teams across Apple to support their workflows or integrate their technology into our platform * Hold yourself and others to a high quality standard expected of Apple products ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) - [How I saved 200K/yr in direct costs writing 0 code lines in K8s](https://www.wearedevelopers.com/videos/1055-how-i-saved-200k-yr-in-direct-costs-writing-0-code-lines-in-k8s) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)