Senior Machine Learning Engineer, Proactive

Apple Inc.
Santa Clara, CA, United States
about 1 month ago
Apply on www.techcareers.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Automated Storage and Retrieval Systems C++ (Programming Language) Encodings Computer Programming Information Retrieval Python (Programming Language) Machine Learning Language Modeling Natural Language Processing Performance Tuning Recommender Systems
+10 more
Tensorflow Search Technologies Pytorch Retrieval-Augmented Generation Large Language Models Deep Learning Generative AI Information Technology Machine Learning Operations GPT

Job description

At Apple, machine learning powers experiences that anticipate what people need before they ask. We’re looking for a Machine Learning Engineer to help build the next generation of intelligent search and AI experiences technology that understands user intent, context, and personal information while preserving privacy. In this role, you’ll design, train, optimize, and deploy large language models, semantic retrieval systems, and ranking models that power relevant, personalized, context-aware search across Apple’s ecosystem. You’ll work at the intersection of search, retrieval, natural language processing, on-device AI, and generative AI to shape the future of intelligent assistants and proactive experiences.., You’ll design, train, fine-tune, and optimize transformer-based language models for on-device deployment, and build semantic retrieval, embedding, reranking, and retrieval-augmented generation systems that improve search quality and AI-powered experiences. You’ll develop models for query understanding, intent prediction, personalization, and ranking, while researching new approaches to model compression, quantization, and low-latency inference. You’ll partner with engineers, researchers, product managers, and designers to bring new AI capabilities from research into production driving technical strategy and leading projects from early exploration through large-scale deployment. This is an opportunity to explore new applications of foundation models, multimodal AI, and agentic retrieval, shaping the next generation of proactive, intelligent user experiences.

Requirements

  • Bachelor’s degree in Computer Science, Machine Learning, Artificial Intelligence, or a related field.
  • 5+ years of industry or research experience developing machine learning systems.
  • Background in machine learning, deep learning, natural language processing, information retrieval, search, recommender systems, or generative AI.
  • Experience training, fine-tuning, or deploying transformer-based models and large language models.
  • Programming skills in Python and/or C/C++, with experience building production-quality software using modern machine learning frameworks such as PyTorch, JAX, or TensorFlow., * Master’s or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, or a related field.
  • Experience optimizing machine learning models for resource-constrained environments, including model compression, quantization, pruning, and knowledge distillation.
  • Experience with on-device machine learning or mobile inference frameworks.
  • Experience building retrieval-augmented generation, vector search, embedding retrieval, or semantic search systems.
  • Experience working with transformer architectures such as BERT, T5, Llama, Gemma, Mistral, or other foundation models.
  • Experience evaluating language models, designing AI quality metrics, and building offline evaluation pipelines.
  • Experience building large-scale production search, recommendation, or personalization systems.
  • Ability to prototype ideas, solve ambiguous problems, and deliver production-quality machine learning solutions.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.techcareers.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · World Congress 2024

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

3:17 min

Optimizing character encoding with Kim variable byte encoding

Douglas Crockford Douglas Crockford · World Congress 2024

40 sec

Navigating Apple's evolving on-device AI and machine learning stack

Precious Osaro Precious Osaro · World Congress 2026 Europe

51 sec

Assessing GPT-4o performance for pull request feedback

Merrill Lutsky Merrill Lutsky · World Congress 2025

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

Videos

See all

Related articles

See all