Machine Learning Architect for Conversational Speech

Apple Inc.
Cupertino, CA, United States
1 day ago
Apply on www.themuse.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$262,500.0
Working hours
Regular working hours

Tech stack

Computer Programming Python (Programming Language) Machine Learning Tensorflow Speech Recognition Pytorch Large Language Models Deep Learning Siri Information Technology Speech Synthesis

Job description

The Speech organization within Siri drives major speech recognition, synthesis, and speech-to-speech model advances for features deeply embedded throughout Apple’s ecosystem. Our mission is to build cutting-edge infrastructure, datasets, and models that empower Siri conversational AI, dictation, and speech-enabled Apple Intelligence features across natural language understanding, dialog generation, speech recognition, and multimodal interaction. We apply these technologies to create engaging, intelligent, and personalized conversational experiences for millions of Apple users., We are seeking a Machine Learning Architect to serve as a senior technical leader spanning the full Speech organization. You will set the future modeling direction for all of conversational speech-charting the architectural and algorithmic course for how Apple’s speech technologies evolve. You will operate as a hands-on expert who not only defines strategy but also digs into the hardest technical problems, working shoulder-to-shoulder with teams to overcome critical obstacles. Reporting directly to the Speech organization leadership, you will have broad visibility and influence across speech recognition, synthesis, dialog, multimodal foundation models, and speech-to-speech systems, ensuring coherent technical vision and cross-team alignment., As the Machine Learning Architect for Conversational Speech, you will define modeling strategy and technical direction across the Speech organization, establishing a unified architectural vision for speech recognition, speech synthesis, dialog systems, multimodal foundation models, and speech-to-speech technologies.

You will serve as the organization’s foremost modeling expert, providing deep technical guidance to multiple teams working on interconnected speech capabilities.

You will evaluate emerging research and industry trends-including advances in large language models, multimodal architectures, and full-duplex natural conversational systems-and translate them into actionable roadmaps.

You will champion production-readiness, ensuring architectural decisions account for on-device constraints, latency, scalability, and robustness.

You will collaborate broadly with partner teams across Siri, Apple Intelligence, hardware, and platform engineering to ensure speech modeling investments are well-integrated into Apple’s broader AI strategy.

Requirements

Ph.D. in Computer Science, Electrical Engineering, Machine Learning, or similar technical field.

Experience architecting or leading development of full-duplex natural conversational systems, speech-to-speech models, or multimodal foundation models that have shipped to large-scale user populations.

Deep familiarity with the full stack of speech technologies-ASR, TTS, spoken dialog, speaker modeling, audio understanding-and an ability to reason about their interactions and dependencies.

Experience with large-scale distributed training and the infrastructure considerations that shape model design at scale.

A data-centric perspective on foundation model development, including experience guiding data collection, curation, annotation, and quality strategies.

Experience with on-device ML deployment, including model compression, quantization, and latency-aware architecture design.

Minimum Qualifications

10+ years of experience in machine learning applied to speech or multimodal systems, with progressively increasing technical scope and leadership.

Demonstrated expertise as a technical leader or architect who has defined modeling direction across multiple teams or product areas., Deep, hands-on proficiency in modern deep learning, including large language models and end-to-end speech systems.

Significant experience with multimodal LLMs, including architecture design, training, adaptation, and deployment of models that integrate speech, audio, and text modalities.

Direct experience building speech-to-speech conversational systems, with a strong understanding of full-duplex natural conversational interaction and end-to-end speech pipelines.

A track record of translating research into production-quality systems at scale.

Expert programming skills in Python and deep learning frameworks such as PyTorch, JAX, or TensorFlow.

Benefits & conditions

At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $262,500 and $394,000, and your base pay will depend on your skills, qualifications, experience, and location.

Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses - including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits

About the company

As part of the Siri organization, you will help shape one of the world’s most widely used AI assistants, powered by our next-generation of Apple Intelligence, with capabilities like personal context understanding and on-screen awareness, built with privacy from the ground up. Your work will have direct, meaningful impact for users across iOS, iPadOS, macOS, watchOS, and visionOS.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.themuse.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:16 min

Addressing latency and architecture in voice agents

Chris Heilmann +2 · LIVE

1:39 min

Fundamentals of tensors and the TensorFlow library

Håkan Silfvernagel · LIVE

1:16 min

Evaluating the enduring financial and technological legacy of Apple

Marco Landi · World Congress 2024

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

3:55 min

Evaluating central server APIs against edge deployment models

Hauke Brammer · World Congress 2021

2:23 min

Historical breakthroughs in natural language processing models

Mary Grygleski Mary Grygleski · LIVE

Videos

See all

Related articles

See all