Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference

Apple Inc.
Seattle, WA, United States
20 days ago
Apply on www.seattlejobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Apple TV App Store (IOS) C++ (Programming Language) Cloud Computing Nvidia CUDA Python (Programming Language) Machine Learning Tensorflow Software Deployment Systems Architecture
+10 more
Private Cloud Environment Pytorch Large Language Models Deep Learning Siri Kubernetes Information Technology Low Latency TensorRT Docker

Job description

We are the Foundation Model Inference team within Cloud OS and AI Inference organization. We are on a mission to build the most highly performant, secure and private inference stack that powers Siri AI, Apple Intelligence and Apps that are powered with the largest foundation models.

Our systems serve billions of queries daily across Siri AI, Apple Intelligence, Apple Search, Apple Music, Apple TV, App Store, iMessage, Photos, Camera, Spotlight & Safari, at remarkably low latency with every ounce of compute extracted from the hardware beneath them. We optimize language, vision, and speech models with billions of parameters using state-of-the-art techniques and ship them at Apple scale.

This is a rare opportunity to directly shape how AI reaches billions of people worldwide., You will work at the intersection of research and production, partnering closely with the Foundation Model Research team and our external partners to bring cutting-edge model architectures from prototype to planetary-scale deployment. You will own hard problems in inference efficiency, hardware/software codesign, systems architecture, and tooling, and help set the technical direction for the engineers around you.

This role sits within CloudOS and Private Cloud Compute (PCC) - Apple’s purpose-built, privacy-preserving cloud infrastructure for AI workloads. PCC represents a first-of-its-kind approach to running foundation models in the cloud with verifiable privacy guarantees, and CloudOS is the systems foundation that makes it possible. You will be building and optimizing inference systems on top of this infrastructure, working closely with platform and security teams to deliver both performance and trust at scale.

Requirements

  • 5+ years of experience leading complex, ambiguous technical projects from end to end.
  • Hands-on experience with LLM inference stacks.
  • Working knowledge of GPU or TPU programming concepts.
  • Proficiency with PyTorch, JAX, or TensorFlow.
  • Experience building and operating high-throughput services at large distributed scale.
  • Proficiency deploying applications on cloud platforms (AWS, GCP, or equivalent) using Kubernetes and Docker.
  • BS in Computer Science, Machine Learning, Artificial Intelligence, Data Science, or a related field., * Experience building productions systems in Go or Python.
  • Strong knowledge of deep learning architectures including Transformers, encoder/decoder models, and multimodal variants.
  • Experience with inference optimization frameworks such as TensorRT-LLM, vLLM, SGLang, TGI, or Nvidia Triton Server.
  • Experience authoring custom CUDA kernels using CUDA C++ or OpenAI Triton.
  • MS in Computer Science, Machine Learning, Artificial Intelligence, Data Science, or a related field.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.seattlejobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · World Congress 2025

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:16 min

Evaluating the enduring financial and technological legacy of Apple

Marco Landi · World Congress 2024

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

2:50 min

Transitioning from deep learning models to foundation software

Marcel Scherenberg Marcel Scherenberg · World Congress 2025

Videos

See all

Related articles

See all