> Markdown version of [/jobs/ext/2306235-software-engineer-platform-reliability-engineering-aidp](https://www.wearedevelopers.com/jobs/ext/2306235-software-engineer-platform-reliability-engineering-aidp). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Platform Reliability Engineering, AiDP - **Company:** Apple Inc. - **Location:** Sunnyvale, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Big Data, Databases, Computer Engineering, Linux, DevOps, Distributed Systems, Python (Programming Language), Machine Learning, Open Source Technology, Systems Architecture, System Programming, Data Processing, System Availability, Apache Spark, HybridCloud, Containerization, AI Platforms, Kubernetes, Information Technology, Apache Flink, Machine Learning Operations, Golang - **Published:** August 30, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3369792056&tx=HT767TYI&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * Bachelor's degree in Computer Science, Computer Engineering, or equivalent professional experience * Proficiency in at least one systems programming language (Python, Go, Java, or similar) * Strong expertise in distributed systems architecture, with deep knowledge of reliability, scalability, and containerization principles * Hands-on experience with cloud platforms and data processing infrastructure (Kubernetes, Spark, Flink, Ray, Trino, or equivalent technologies), * 7+ years of experience in SRE, DevOps, or infrastructure engineering, with demonstrated expertise managing distributed systems at scale. * Proficiency in diagnosing and resolving complex production incidents and performance bottlenecks in large-scale distributed environments. * Familiarity with open source codebases; ability to read, understand, and explain complex system implementations * Strong understanding of system architecture and proven ability to collaborate effectively across engineering teams * Hands-on experience with big data technologies (Spark, Flink, Iceberg) and/or ML/AI platforms (Ray, MLflow, model serving infrastructure). * Strong foundational knowledge of Linux, databases, and security principles * Proactive mindset with demonstrated commitment to optimizing reliability and uptime for mission-critical services * Excellent written and verbal communication skills with ability to articulate technical concepts and strategies to both engineering teams and non-technical leadership * Demonstrated track record of designing and operating systems at scale ## Description We're seeking an experienced software engineer to join our Platform Reliability Engineering team and drive the design, operation, and optimization of large-scale distributed systems that power our GenAI, ML, and big data platforms. You'll leverage cutting-edge open source technologies in hybrid cloud environments to build resilient infrastructure that enables seamless inference, data processing, and machine learning workloads at scale. In this role, you'll own mission-critical platform components, respond to production incidents, and collaborate across teams to shape the future of our data and AI infrastructure. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Platform Engineering vs. DevOps Why not both?](https://www.wearedevelopers.com/videos/885-platform-engineering-vs-devops-why-not-both) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)