> Markdown version of [/jobs/ext/2402482-multimodal-ai-researcher](https://www.wearedevelopers.com/jobs/ext/2402482-multimodal-ai-researcher). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Multimodal AI Researcher - **Company:** Apple Inc. - **Location:** Sunnyvale, CA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computer Vision, Computer Engineering, Computer Graphics, Python (Programming Language), Machine Learning, Software Engineering, Strategies of Testing, Reinforcement Learning, Pytorch, Large Language Models, Facebook Flow, Information Technology, Stable Diffusion - **Published:** August 13, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3351752107&tx=FR5754FFL&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role We hire researchers who are highly motivated and deeply care about shipping. A successful candidate will stay up-to-date with the latest advancements in multimodal foundations models and applying this knowledge to drive innovation, but also take a practical approach to problem solving and software engineering., * BS and a minimum of 3 years relevant industry experience. * Experience building models for multimodal perception systems. * Experience working with LLMs and VLMs. * Software engineering skills and proficiency in Python and PyTorch. * Curiosity and willingness to learn new things in order to improve the quality of their solutions., * MS or PhD in computer vision, computer graphics, machine learning, computer science, computer engineering or related fields. * Experience in developing, training/tuning foundation models and multimodal LLMs. * Experience with training and troubleshooting generative architectures such as diffusion, reinforcement learning, flow matching or normalizing flow at scale. * Experience with real-time or streaming multimodal models. * Experience with speech understanding and generation. * Experience applying reinforcement learning to help post-train foundation models. * Excellent communication and experience working with multi-functional teams. * Self-motivated with proven track record to optimally prioritize and deliver tasks on schedule. ## Description We are looking for a Multimodal AI Researcher with a strong background in developing foundation models for generative AI and multimodal systems that integrate various types of real-time sensor data such as video and audio with other modalities like text. Our ongoing investigations include interactive models, audio-to-audio modeling and systems. You will work on hard, open research problems in multimodal generative AI and agents, and you will see that work through to real features used by millions of people. You will collaborate with others to drive data requirements, validation strategies, and key performance indicators, and conduct algorithm research and development that serves product needs. ## Related Videos - [Multimodal Generative AI Demystified](https://www.wearedevelopers.com/videos/829-multimodal-generative-ai-demystified) - [Focoos AI: Building the Future of Computer Vision](https://www.wearedevelopers.com/videos/1659-focoos-ai-building-the-future-of-computer-vision) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Edge AI on iOS: Beyond the Cloud, Designing the Next Generation of Intelligent On-Device Apps](https://www.wearedevelopers.com/videos/100225-edge-ai-on-ios-beyond-the-cloud-designing-the-next-generation-of-intelligent-on-device-apps) - [Computer Vision from the Edge to the Cloud done easy](https://www.wearedevelopers.com/videos/263-computer-vision-from-the-edge-to-the-cloud-done-easy) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)