Machine Learning Engineer, Foundation Model Services

Apple Inc.
Seattle, WA, United States
22 days ago
Apply on www.seattlejobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Regular working hours

Tech stack

Clean Code Principles Amazon Web Services App Store (IOS) Microsoft Azure Information Retrieval Python (Programming Language) Machine Learning Natural Language Processing Google Cloud Cloud Platform System Large Language Models Siri
+8 more
Kubernetes Information Technology Low Latency Build Tools Machine Learning Operations TensorRT Docker Programming Languages

Job description

Do you think differently? Are you eager to break the status quo, bold and ambitious, unafraid to take risks, and passionate about building best-in-class technology? If so, there’s no better place to do it than Apple. The Foundation Model Services team builds the frameworks, services, and tools that run Apple’s largest foundation models in production. Our infrastructure powers intelligent experiences across products people use every day - Search, Music, TV, the App Store, Messages, Photos, Spotlight, Safari, Siri, and more - serving millions of queries at incredibly low latency while drawing every ounce of performance from our hardware. Join us and you’ll help bring intelligence to billions of users around the world, working on optimizing and serving large language, vision, and speech models at Apple’s scale., Work closely with product teams to build production-grade solutions that launch models serving customers in real time. Partner with foundation model researchers to prototype and develop inference for cutting-edge model architectures, and build tools that help us understand and remove performance bottlenecks across different hardware and use cases. Write high-quality code, learn quickly in a fast-moving field, and grow your impact as you take on larger pieces of the system.

Requirements

  • 2+ years of industry experience building and shipping production software and/or machine learning systems.
  • Proficiency in a modern programming language such as Go or Python.
  • Experience deploying and operating services on a cloud platform (AWS, Azure, GCP, or equivalent) using containers and Kubernetes/Docker.
  • 5 year+ industry experience in ML technologies (LLMs, Machine Learning, NLP, Information Retrieval, Statistics).
  • Experience building or operating high-throughput, low-latency services.
  • Strong communication and collaboration skills, with the ability to partner across research and product teams.
  • Bachelor’s degree or higher in Computer Science or related technical field.

Preferred Qualifications

  • Familiarity with Nvidia TensorRT-LLM, vLLLM, DeepSpeed, Nvidia Triton Server etc.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.seattlejobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:16 min

Evaluating the enduring financial and technological legacy of Apple

Marco Landi · World Congress 2024

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · World Congress 2025

2:50 min

Transitioning from deep learning models to foundation software

Marcel Scherenberg Marcel Scherenberg · World Congress 2025

2:23 min

Historical breakthroughs in natural language processing models

Mary Grygleski Mary Grygleski · LIVE

1:12 min

Training and fine-tuning models natively using MLX

MIlan Todorović MIlan Todorović · World Congress 2025

Videos

See all

Related articles

See all