> Markdown version of [/jobs/ext/1628971-research-engineer-computer-vision](https://www.wearedevelopers.com/jobs/ext/1628971-research-engineer-computer-vision). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Research Engineer, Computer Vision - **Company:** Facebook Inc. - **Location:** Pittsburgh, PA, United States - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Algorithm Design, Computer Vision, C++ (Programming Language), Python (Programming Language), Machine Learning, Language Modeling, Tensorflow, Pytorch, Large Language Models, Deep Learning, Information Technology - **Published:** July 1, 2026 - **Apply:** https://www.jobmonkeyjobs.com/career/27805705/Research-Engineer-Computer-Vision-Pennsylvania-Pittsburgh-7418 ## About the Role * Proven experience with C++ and/or Python, including experience with modern features * Experience working with deep learning frameworks such as PyTorch and TensorFlow * Demonstrated experience working collaboratively in cross-functional teams Preferred Qualifications: * Master's degree in Computer Science, Computer Vision, Machine Learning, or related field * Experience with vision-language models or multi-modal transformers * Publications or contributions to multi-modal understanding research * Familiarity with large language models and their integration with visual understanding systems * Experience with data curation, annotation tools, or ground truth labeling pipelines ## Description As a Research Engineer focused on Multi-Modal Understanding, you will develop advanced algorithms that integrate computer vision with other modalities such as language, audio, and sensor data. You will also drive the curation of multi-modal datasets and ground truth annotation pipelines to support model training and evaluation. You will work closely with our research team to bring innovative multi-modal solutions to production, bridging the gap between visual perception and holistic contextual understanding for immersive applications., * Design and implement multi-modal understanding systems that combine vision, language, and other sensory inputs to enable richer contextual awareness * Develop algorithms for cross-modal learning, fusion, and reasoning to improve human-AI interaction * Lead the curation and management of multi-modal datasets, ensuring data quality and diversity across vision, language, and sensor modalities * Design and oversee ground truth annotation workflows and quality assurance processes for multi-modal data * Complete medium to large features spanning multiple tasks independently with minimal to no guidance * Collaborate with researchers and engineers across computer vision and machine learning teams to drive multi-modal innovation * Develop well-organized code with proper testing and documentation, building production-ready multi-modal systems ## Related Videos - [Getting Started with Machine Learning](https://www.wearedevelopers.com/videos/260-getting-started-with-machine-learning) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Focoos AI: Building the Future of Computer Vision](https://www.wearedevelopers.com/videos/1659-focoos-ai-building-the-future-of-computer-vision) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [30 Golden Rules of Deep Learning Performance](https://www.wearedevelopers.com/videos/11-30-golden-rules-of-deep-learning-performance) ## Related Articles - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [What’s in The box? – Unboxing The DeepFace](https://www.wearedevelopers.com/magazine/117-what-s-in-the-box-unboxing-the-deepface) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [How to start an AI project for a good cause and boost your career](https://www.wearedevelopers.com/magazine/15-how-to-start-an-ai-project-for-a-good-cause-and-boost-your-career) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)