> Markdown version of [/jobs/ext/2833307-site-reliability-engineer-ml-apple-ads](https://www.wearedevelopers.com/jobs/ext/2833307-site-reliability-engineer-ml-apple-ads). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - ML, Apple Ads - **Company:** Apple Inc. - **Location:** New York, NY, United States - **Experience:** Experienced - **Salary:** $150,400.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Airflow, Amazon Web Services, Computer Programming, Continuous Integration, Linux, DevOps, Distributed Systems, Python (Programming Language), Machine Learning, Network Architecture, Reliability Engineering, Azure Machine Learning, Rust (Programming Language), Reliability of Systems, Backend, Kubernetes, Machine Learning Operations, Terraform, Golang - **Published:** September 10, 2026 - **Apply:** https://www.jobmonkeyjobs.com/career/28006420/Site-Reliability-Engineer-Ml-Apple-Ads-New-York-New-York-7413 ## About the Role 3+ years of experience in internet-facing backend production systems, SRE or ML Operations focused roles on large scale distributed cloud infrastructure Proven expertise with AWS-managed infrastructure Familiarity with ML lifecycle and associated technologies such as NVIDIA Triton, AnyScale Ray, Apache Airflow etc. Strong programming skills in at least one of: Python, Java, Rust, Go or similar languages Hands-on experience with Linux systems and deep knowledge of its internals. Demonstrated experience with Infrastructure as Code, especially Terraform. Strong foundation in SRE concepts: Monitoring, alerting, observability, Incident response and root cause analysis, Error budgets, SLAs/SLOs, and system reliability Preferred Qualifications Built tools or services that automate platform operations, reduce toil, or improve cost efficiency. Experience managing Kubernetes clusters at scale in production environments. Hands-on experience troubleshooting distributed systems under real-world load. Clear communication skills and comfort collaborating across engineering, infrastructure, and product teams. AWS certifications or broad experience across multiple AWS services is a plus. Understanding of modern GPU hardware architectures (such as NVIDIA H100, B200, or GB200, AWS Inferentia ), associated drivers Understanding of high-performance fabrics and network architecture, power, and thermal limits ## Description We are looking for a ML Platform Infrastructure Engineer to help build and evolve the next generation of Apple Ads machine learning platform - enabling fast, reliable, and scalable operations across AWS-based environments supporting transactional and analytical workloads., As a site reliability engineer in Apple Ads focused on machine learning, you will own the health, performance, and scalability of large scale infrastructure powering ML training, inference, serving workloads and associated platform tooling. Your focus will be on building automation that eliminates manual processes, improves platform resilience, and enables teams to move faster with confidence. This is not a DevOps-only or CI/CD-focused role. We are looking for engineers who build platform solutions, not just configure pipelines. Responsibilities Build and operate distributed systems using AWS managed services such as EKS, ElasticCache and ML technologies like Ray over Kubernetes and NVIDIA Triton Inference Server. Develop internal tooling and automation frameworks to improve infrastructure reliability, cost-efficiency, and operational visibility. Collaborate with engineering teams to define infrastructure architecture, troubleshoot complex issues, and drive production excellence. Design and manage Infrastructure as Code with Terraform, ensuring repeatable, secure, and scalable deployments. Lead or participate in incident response, postmortems, and continuous improvement cycles to reduce future risk. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [7 Most Popular Web Developer Jobs in Europe](https://www.wearedevelopers.com/magazine/163-7-most-popular-web-developer-jobs-in-europe) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)