> Markdown version of [/jobs/ext/2697213-machine-learning-research-engineer-multimodal-for-human-understanding](https://www.wearedevelopers.com/jobs/ext/2697213-machine-learning-research-engineer-multimodal-for-human-understanding). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Research Engineer - Multimodal for Human Understanding - **Company:** Apple Inc. - **Location:** Sunnyvale, CA, United States - **Experience:** Experienced - **Salary:** $150,400.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Apple Products, Computer Vision, C++ (Programming Language), Python (Programming Language), Machine Learning, Reinforcement Learning, Pytorch, Large Language Models, Deep Learning, Generative AI, Information Technology, Machine Learning Operations, Virtual Agents - **Published:** September 3, 2026 - **Apply:** https://www.themuse.com/jobs/apple/applied-machine-learning-research-engineer-multimodal-for-human-understanding-333682 ## About the Role Hands-on experience training and deploying production-grade ML models. Experience developing multimodal LLMs or generative models. Production-level experience with a compiled language (e.g., Swift, C++). Expertise in one or more areas: computer vision, machine learning, multimodal LLMs, Reinforcement Learning, Agentic AI. PhD in Computer Science, Electrical Engineering, or a related field with a focus on computer vision, machine learning, or multimodal systems. Demonstrated problem-solving ability, strong sense of ownership and product shipment. Minimum Qualifications Strong experience developing machine learning models. Proficiency in Python and solid software engineering fundamentals. Experience with at least one deep learning framework (e.g., PyTorch, JAX, or equivalent). Master's degree in Computer Science or a related field, plus 3 years of relevant industry experience. ## Description We're starting to see the incredible potential of multimodal foundation and large language models, and many applications in the computer vision and machine learning domain that previously appeared infeasible are now within reach. We are looking for a highly motivated and skilled Applied Machine Learning Research Engineer to join our team in the Video Computer Vision group and help us push the boundaries of human understanding. The Video Computer Vision org has pioneered human-centric real-time features such as FaceID, FaceKit, and Gaze and Hand gesture control which have changed the way millions of users interact with their devices. We balance research and product requirements to deliver Apple quality, pioneering experiences, innovating through the full stack, and partnering with HW, SW and AI teams to shape Apple's products and bring our vision to life., In this role, you will drive ground breaking development at the intersection of AI, generative modeling, and computer vision. You will work across the full lifecycle-from foundational investigation to practical applications-designing, implementing, and evaluating novel algorithms and models. Your primary focus will be human understanding, including human motion, activities, and representation learning. A major aspect of the role involves designing, implementing, evaluating and productizing ML systems capable of human and activity understanding. This position offers a unique opportunity to innovate, build, and ship: you will take your conceptual ideas to products that reach millions of users worldwide. You will collaborate with a diverse group of experts-research scientists, ML engineers, software engineers, data scientists, human-interface designers, and domain specialists-working in an environment that values experimentation, ownership, and continuous learning. By staying at the forefront of advancements in AI, machine learning, and computer vision, you will play a direct role in driving innovation, influencing the evolution of Apple products, and meaningfully enhancing user experience on a global scale. ## Related Videos - [Getting Started with Machine Learning](https://www.wearedevelopers.com/videos/260-getting-started-with-machine-learning) - [Multimodal Generative AI Demystified](https://www.wearedevelopers.com/videos/829-multimodal-generative-ai-demystified) - [Your imaginations is (no longer) the limit: how Generative AI empowers people to be creative](https://www.wearedevelopers.com/videos/741-your-imaginations-is-no-longer-the-limit-how-generative-ai-empowers-people-to-be-creative) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Inside the Mind of an LLM](https://www.wearedevelopers.com/videos/1617-inside-the-mind-of-an-llm) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How machine learning can help us tell fact from fiction](https://www.wearedevelopers.com/magazine/509-how-machine-learning-can-help-us-tell-fact-from-fiction) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)