Multimodal AI Researcher

Apple Inc.
Sunnyvale, CA, United States
23 days ago
Apply on www.techcareers.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Computer Vision Computer Engineering Computer Graphics Python (Programming Language) Machine Learning Software Engineering Strategies of Testing Reinforcement Learning Pytorch Large Language Models Facebook Flow
+2 more
Information Technology Stable Diffusion

Job description

We are looking for a Multimodal AI Researcher with a strong background in developing foundation models for generative AI and multimodal systems that integrate various types of real-time sensor data such as video and audio with other modalities like text. Our ongoing investigations include interactive models, audio-to-audio modeling and systems. You will work on hard, open research problems in multimodal generative AI and agents, and you will see that work through to real features used by millions of people. You will collaborate with others to drive data requirements, validation strategies, and key performance indicators, and conduct algorithm research and development that serves product needs.

Requirements

We hire researchers who are highly motivated and deeply care about shipping. A successful candidate will stay up-to-date with the latest advancements in multimodal foundations models and applying this knowledge to drive innovation, but also take a practical approach to problem solving and software engineering., * BS and a minimum of 3 years relevant industry experience.

  • Experience building models for multimodal perception systems.
  • Experience working with LLMs and VLMs.
  • Software engineering skills and proficiency in Python and PyTorch.
  • Curiosity and willingness to learn new things in order to improve the quality of their solutions., * MS or PhD in computer vision, computer graphics, machine learning, computer science, computer engineering or related fields.
  • Experience in developing, training/tuning foundation models and multimodal LLMs.
  • Experience with training and troubleshooting generative architectures such as diffusion, reinforcement learning, flow matching or normalizing flow at scale.
  • Experience with real-time or streaming multimodal models.
  • Experience with speech understanding and generation.
  • Experience applying reinforcement learning to help post-train foundation models.
  • Excellent communication and experience working with multi-functional teams.
  • Self-motivated with proven track record to optimally prioritize and deliver tasks on schedule.

About the company

The Video Computer Vision organization is working on breakthrough technologies for future Apple products. Our team delivers cutting-edge AI, machine learning, computer vision and graphics algorithms that power technologies including human understanding, perception, digital humans, multimodal generative AI, and agents. Our algorithms ship across a range of Apple products, including iPhone and Apple Vision Pro, where our work has contributed to technologies like Personalized Spatial Audio, EyeSight, and Persona as well as future Apple products. We are an applied research group, we push the state of the art and then bring it to product. In this role, you will collaborate with world-class experts in AI, ML, Software, and Hardware to tackle fundamental challenges in human-centric solutions that will impact millions of users across Apple’s ecosystem.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.techcareers.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:55 min

Practical applications and use cases for computer vision

Flo Pachinger · LIVE

1:41 min

Focusing on core functional features instead of cosmetic details

Kristijan Ristovski · World Congress 2022

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

3:22 min

Evaluating advanced artificial intelligence platforms for daily recruitment

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

51 sec

Overcoming inefficiencies in computer vision modeling

Antonio Tavera Antonio Tavera · World Congress 2025

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all