Senior Machine Learning Engineer, Proactive

Apple Inc.
Santa Clara, CA, United States
17 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$184,700.0 - $324,800.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Automated Storage and Retrieval Systems C++ (Programming Language) Encodings Computer Programming Information Retrieval Python (Programming Language) Machine Learning Language Modeling Natural Language Processing Performance Tuning Recommender Systems
+12 more
Tensorflow Search Technologies Pytorch Retrieval-Augmented Generation Large Language Models Deep Learning Generative AI Information Technology Low Latency Machine Learning Operations Virtual Agents GPT

Job description

You’ll design, train, fine-tune, and optimize transformer-based language models and foundation models for efficient on-device deployment, and build semantic retrieval, embedding, reranking, and retrieval-augmented generation systems that improve search quality and AI-powered experiences. You’ll develop models for query understanding, intent prediction, personalization, retrieval, and ranking, while researching new approaches to LLM fine-tuning, knowledge distillation, model compression, quantization, and low-latency inference. You’ll explore techniques for adapting large foundation models into smaller, highly capable models that can operate efficiently under on-device memory, compute, power, and latency constraints.

You’ll partner with engineers, researchers, product managers, and designers to bring new AI capabilities from research into production, driving technical strategy and leading projects from early exploration through large-scale deployment. This is an opportunity to explore new applications of foundation models, multimodal AI, agentic retrieval, and personalized intelligence, shaping the next generation of proactive and intelligent user experiences.”,”responsibilities”:”Build semantic retrieval, embedding, reranking, and retrieval-augmented generation systems, along with models for query understanding, intent prediction, personalization, retrieval, and ranking.

Analyze search relevance and user behavior to design evaluation methodologies, offline benchmarks, and online metrics that measure retrieval quality, ranking, personalization, and language model performance.

Build scalable experimentation and evaluation pipelines for LLMs and search models, including model quality, robustness, latency, efficiency, and end-to-end product metrics.

Design, train, fine-tune, distill, and optimize transformer-based language models and foundation models for efficient on-device deployment.

Develop LLM fine-tuning and post-training approaches, including supervised fine-tuning, instruction tuning, preference optimization, parameter-efficient fine-tuning, and task-specific adaptation.

Research and prototype approaches for on-device generative AI, including knowledge distillation, model compression, quantization, pruning, and low-latency inference.

Develop techniques to transfer capabilities from large foundation models into compact on-device models while balancing model quality, latency, memory footprint, power consumption, and compute constraints.

Partner with engineers, researchers, product managers, and designers to bring AI capabilities from research into production, driving technical strategy across projects and exploring new applications of foundation models, multimodal AI, agentic retrieval, and personalized intelligence.

Requirements

Master’s or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, or a related field.

Experience optimizing machine learning models for resource-constrained environments, including knowledge distillation, model compression, quantization, and pruning.

Experience with on-device machine learning or edge AI, or mobile inference frameworks, including optimizing models for latency, memory, compute, and power constraints.

Experience distilling capabilities from large foundation models into small language models or task-specific models for efficient inference.

Experience building retrieval-augmented generation, vector search, embedding retrieval, neural reranking, or semantic search systems.

Experience with query understanding, query rewriting, intent classification, personalized retrieval, learning-to-rank, or recommendation models.

Experience working with transformer architectures and foundation model families such as BERT, T5, Llama, Gemma, Mistral, or related architectures.

Experience evaluating language models, designing AI quality metrics, and building automated and human-in-the-loop evaluation pipelines.

Experience building large-scale production search, recommendation, personalization, or generative AI systems.

Familiarity with multimodal foundation models, tool use, agentic AI, or agentic retrieval systems.

Strong understanding of the tradeoffs among model quality, latency, memory, power consumption, privacy, and reliability for production on-device AI systems.

Ability to prototype new ideas, conduct rigorous experiments, solve ambiguous technical problems, and translate research advances into production-quality machine learning solutions.

Minimum Qualifications

Master degree in Computer Science, Machine Learning, Artificial Intelligence, or a related field.

5+ years of industry or research experience developing machine learning systems.

Background in machine learning, deep learning, natural language processing, information retrieval, search, recommender systems, or generative AI.

Experience training, fine-tuning, or deploying transformer-based models and large language models.

Experience with modern deep learning architectures and techniques, including transformers, embeddings, representation learning, and neural ranking.

Programming skills in Python and/or C/C++, with experience building production-quality software using modern machine learning frameworks such as PyTorch, JAX, or TensorFlow.

Ability to work onsite in Cupertino, California, in accordance with Apple’s applicable work policies.

Benefits & conditions

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $184,700 and $324,800, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses - including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

40 sec

Navigating Apple's evolving on-device AI and machine learning stack

Precious Osaro Precious Osaro · World Congress 2026 Europe

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · World Congress 2024

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

3:17 min

Optimizing character encoding with Kim variable byte encoding

Douglas Crockford Douglas Crockford · World Congress 2024

2:01 min

Exploring foundational expertise in traditional optimization and machine learning

Eric Enge · Coffee With Developers

51 sec

Assessing GPT-4o performance for pull request feedback

Merrill Lutsky Merrill Lutsky · World Congress 2025

Videos

See all

Related articles

See all