> Markdown version of [/jobs/ext/735439-ai-researcher-computer-vision-multimodal-generative-ai](https://www.wearedevelopers.com/jobs/ext/735439-ai-researcher-computer-vision-multimodal-generative-ai). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Researcher (Computer Vision/Multimodal/Generative AI) - **Company:** SpreeAI Corporation - **Location:** San Francisco, CA, United States - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computer Vision, Computer Programming, High-Level Architecture, Python (Programming Language), Machine Learning, Object-Oriented Software Development, Pytorch, Deep Learning, Generative AI, Information Technology, GPT - **Published:** June 29, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=91c953ea8bd76315 ## About the Role Do you have experience in Research design?, * PhD in Computer Science, Artificial Intelligence, Robotics, Computer Vision, or related field. * Strong research background in computer vision, generative modeling, or multimodal AI. * Strong programming skills in Python and familiarity with object-oriented languages. * Experience with deep learning frameworks (PyTorch preferred). * Strong foundations in machine learning theory and experimental design. Preferred Qualifications * Publications at top conferences (CVPR, ICCV, NeurIPS, ICLR, SIGGRAPH, etc.). * Experience with diffusion-based generative models. * Video modeling or temporal learning experience. * Experience bridging research into production systems. * Interest in compute efficiency, distillation, or scalable generative pipelines. ## Description We are hiring ML Researchers to develop novel approaches that advance the frontier of multimodal vision AI and create product-defining capabilities for SpreeAI. This role exists because current generative and vision models are not designed for photorealistic human representation, controllable try-on, or real-world deployment constraints. You will explore new architectures, algorithms, and training strategies that improve realism, controllability, efficiency, and multimodal understanding - with a direct path from research to production. You will work on research problems across: * photorealistic virtual try-on * human-centric visual representation learning * video-based modeling and temporal consistency * multimodal reasoning and generative pipelines * compute-efficient diffusion and generative architectures This is a research role with product impact: successful work leads to platform capabilities, white papers, patents and most importantly, industry differentiation. Why This Role Exists Modern multimodal AI systems struggle with identity preservation, pose consistency, physical realism, and controllability under production constraints. We are building new approaches where: * diffusion models must produce consistent outputs across poses, viewpoints, and garments, * generative models must learn human and garment interactions realistically, * research innovations must scale to real-world deployment environments. This role is for researchers who want to see novel ideas become shipped systems used by real customers. What you'll do * Develop novel architectures and training approaches for vision and multimodal AI. * Advance generative modeling techniques including controllable diffusion and video generation. * Design experiments improving realism, temporal consistency, and human representation. * Collaborate with applied engineering teams to translate research into production systems. * Publish white papers or research outputs aligned with product differentiation. * Evaluate new model paradigms for scalability and efficiency. Core Research Areas & Model Architectures Candidates should have familiarity with or interest in advancing: * Diffusion models and latent diffusion architectures. * Transformer-based vision models (ViT, multimodal transformers). * Image-to-image and video generation pipelines. * Control mechanisms for generative models (conditioning, adapters, LoRA). * Representation learning for human pose, geometry, or identity consistency. * Multimodal architectures combining vision, text, and structured inputs. ## Related Videos - [Getting Started with Machine Learning](https://www.wearedevelopers.com/videos/260-getting-started-with-machine-learning) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Should we build Generative AI into our existing software?](https://www.wearedevelopers.com/videos/1129-should-we-build-generative-ai-into-our-existing-software) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Speak, Code, Deploy: Transforming Developer Experience with Voice Commands](https://www.wearedevelopers.com/videos/1159-speak-code-deploy-transforming-developer-experience-with-voice-commands) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [DeepSeek R1 vs ChatGPT o1: How Do They Compare?](https://www.wearedevelopers.com/magazine/542-deepseek-r1-vs-chatgpt-o1-how-do-they-compare) - [How to start an AI project for a good cause and boost your career](https://www.wearedevelopers.com/magazine/15-how-to-start-an-ai-project-for-a-good-cause-and-boost-your-career) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)