Site Reliability Engineer - ML, Apple Ads

Apple Inc.
New York, NY, United States
11 days ago
Apply on www.jobmonkeyjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$150,400.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Airflow Amazon Web Services Computer Programming Continuous Integration Linux DevOps Distributed Systems Python (Programming Language) Machine Learning Network Architecture Reliability Engineering
+8 more
Azure Machine Learning Rust (Programming Language) Reliability of Systems Backend Kubernetes Machine Learning Operations Terraform Golang

Job description

We are looking for a ML Platform Infrastructure Engineer to help build and evolve the next generation of Apple Ads machine learning platform - enabling fast, reliable, and scalable operations across AWS-based environments supporting transactional and analytical workloads., As a site reliability engineer in Apple Ads focused on machine learning, you will own the health, performance, and scalability of large scale infrastructure powering ML training, inference, serving workloads and associated platform tooling. Your focus will be on building automation that eliminates manual processes, improves platform resilience, and enables teams to move faster with confidence.

This is not a DevOps-only or CI/CD-focused role. We are looking for engineers who build platform solutions, not just configure pipelines.

Responsibilities

Build and operate distributed systems using AWS managed services such as EKS, ElasticCache and ML technologies like Ray over Kubernetes and NVIDIA Triton Inference Server.

Develop internal tooling and automation frameworks to improve infrastructure reliability, cost-efficiency, and operational visibility.

Collaborate with engineering teams to define infrastructure architecture, troubleshoot complex issues, and drive production excellence.

Design and manage Infrastructure as Code with Terraform, ensuring repeatable, secure, and scalable deployments.

Lead or participate in incident response, postmortems, and continuous improvement cycles to reduce future risk.

Requirements

3+ years of experience in internet-facing backend production systems, SRE or ML Operations focused roles on large scale distributed cloud infrastructure

Proven expertise with AWS-managed infrastructure

Familiarity with ML lifecycle and associated technologies such as NVIDIA Triton, AnyScale Ray, Apache Airflow etc.

Strong programming skills in at least one of: Python, Java, Rust, Go or similar languages

Hands-on experience with Linux systems and deep knowledge of its internals.

Demonstrated experience with Infrastructure as Code, especially Terraform.

Strong foundation in SRE concepts: Monitoring, alerting, observability, Incident response and root cause analysis, Error budgets, SLAs/SLOs, and system reliability

Preferred Qualifications

Built tools or services that automate platform operations, reduce toil, or improve cost efficiency.

Experience managing Kubernetes clusters at scale in production environments.

Hands-on experience troubleshooting distributed systems under real-world load.

Clear communication skills and comfort collaborating across engineering, infrastructure, and product teams.

AWS certifications or broad experience across multiple AWS services is a plus.

Understanding of modern GPU hardware architectures (such as NVIDIA H100, B200, or GB200, AWS Inferentia ), associated drivers

Understanding of high-performance fabrics and network architecture, power, and thermal limits

Benefits & conditions

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $150,400 and $225,300, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses - including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits

Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

About the company

At Apple, we focus deeply on our customers’ experience. Apple Ads brings this same approach to advertising, helping people find exactly what they’re looking for and helping advertisers grow their businesses.

Our technology powers ads and sponsorships across Apple Services, including the App Store, Apple News, and MLS Season Pass. Everything we do is designed for trust, connection, and impact: We respect user privacy, integrate advertising thoughtfully into the experience, and deliver value for advertisers of all sizes-from small app developers to big, global brands. Because when advertising is done right, it benefits everyone.

The Site Reliability Engineering team within Apple Ads ensures the reliability, performance, and availability of ML Platform and Services at scale. The team partners closely with Ads engineering, data science and ML platform teams to enable product delivery through design, configuration, and automation of machine learning infrastructure powering Apple Ads applications.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jobmonkeyjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:38 min

Adopting site reliability engineering practices for machine learning

Cassie Kozyrkov · World Congress 2022

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all