Software Engineer, Platform Reliability Engineering, AiDP

Apple Inc.
Sunnyvale, CA, United States
3 days ago
Apply on www.techcareers.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours

Tech stack

Java (Programming Language) Big Data Databases Computer Engineering Linux DevOps Distributed Systems Python (Programming Language) Machine Learning Open Source Technology Systems Architecture System Programming
+11 more
Data Processing System Availability Apache Spark HybridCloud Containerization AI Platforms Kubernetes Information Technology Apache Flink Machine Learning Operations Golang

Job description

We’re seeking an experienced software engineer to join our Platform Reliability Engineering team and drive the design, operation, and optimization of large-scale distributed systems that power our GenAI, ML, and big data platforms. You’ll leverage cutting-edge open source technologies in hybrid cloud environments to build resilient infrastructure that enables seamless inference, data processing, and machine learning workloads at scale. In this role, you’ll own mission-critical platform components, respond to production incidents, and collaborate across teams to shape the future of our data and AI infrastructure.

Requirements

  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent professional experience
  • Proficiency in at least one systems programming language (Python, Go, Java, or similar)
  • Strong expertise in distributed systems architecture, with deep knowledge of reliability, scalability, and containerization principles
  • Hands-on experience with cloud platforms and data processing infrastructure (Kubernetes, Spark, Flink, Ray, Trino, or equivalent technologies), * 7+ years of experience in SRE, DevOps, or infrastructure engineering, with demonstrated expertise managing distributed systems at scale.
  • Proficiency in diagnosing and resolving complex production incidents and performance bottlenecks in large-scale distributed environments.
  • Familiarity with open source codebases; ability to read, understand, and explain complex system implementations
  • Strong understanding of system architecture and proven ability to collaborate effectively across engineering teams
  • Hands-on experience with big data technologies (Spark, Flink, Iceberg) and/or ML/AI platforms (Ray, MLflow, model serving infrastructure).
  • Strong foundational knowledge of Linux, databases, and security principles
  • Proactive mindset with demonstrated commitment to optimizing reliability and uptime for mission-critical services
  • Excellent written and verbal communication skills with ability to articulate technical concepts and strategies to both engineering teams and non-technical leadership
  • Demonstrated track record of designing and operating systems at scale

About the company

AI & Data Platforms (AiDP) is IS&T’s engine for AI-powered innovation. The team brings together data, application development, and machine learning - including generative AI - along with data services and customer success functions, to help IS&T build solutions more efficiently and streamline the adoption and embedding of generative AI across Apple.

The Applied Machine Learning team in AI and Data Platform organization is building the foundation for Apple’s enterprise-wide machine learning and data capabilities. Our Applied Machine Learning team designs, builds, and operates mission-critical platforms and services spanning ML, GenAI, inference, and big data-enabling teams across the company to harness AI and analytics at scale. We tackle complex technical challenges in reliability, performance, and scalability across a diverse ecosystem of open source and cutting-edge technologies, serving some of Apple’s most demanding workloads.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.techcareers.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · World Congress 2021

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all