> Markdown version of [/jobs/ext/2720679-vision-language-model-engineer](https://www.wearedevelopers.com/jobs/ext/2720679-vision-language-model-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Vision Language Model Engineer - **Company:** EchoTwin AI, Inc. - **Location:** San Francisco, CA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Artificial Neural Networks, Computer Vision, Microsoft Azure, Cloud Computing, Data Transformation, Distributed Computing Environment, Python (Programming Language), Machine Learning, Language Modeling, Natural Language Processing, OpenCV, Tensorflow, Google Cloud, Pytorch, Large Language Models, Deep Learning, Generative AI, Question Answering, Scikit Learn, Information Technology, Low Latency, HuggingFace, Machine Learning Operations - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/vision-language-model-engineer-echotwin-ai-8116321 ## About the Role * Bachelor's, Master's or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, or a related field (or equivalent experience). * 3+ years of experience in machine learning, with a focus on vision-language models or multimodal AI. * Hands-on experience with deep learning frameworks such as PyTorch or TensorFlow. * Proven track record of building and deploying computer vision and/or NLP models. * Proficiency in Python and relevant ML libraries (e.g., Hugging Face, OpenCV, Transformers). * Experience with large-scale model training and optimization (e.g., distributed training, quantization). * Strong understanding of neural network architectures (e.g., CNNs, Transformers, CLIP, or similar). * Experience with multimodal datasets and preprocessing techniques for images and text. * Familiarity with cloud platforms (e.g., AWS, GCP, Azure) and model deployment workflows. * Strong problem-solving skills and ability to work in a fast-paced, collaborative environment. * Excellent communication skills to explain complex technical concepts to diverse audiences. ## Description As a Vision Language Model Engineer, you will design, develop, and optimize advanced vision-language models that integrate visual and textual data to enable intelligent systems. You will work closely with cross-functional teams to build models that power applications such as image captioning, visual question answering, and multimodal AI at the edge., * Design and implement state-of-the-art vision-language models using deep learning frameworks. * Develop and fine-tune models that combine computer vision and natural language processing for tasks like image captioning, visual question answering, and text-to-image generation. * Collaborate with data scientists and software engineers to integrate models into production systems. * Optimize model performance for accuracy, latency, and scalability in real-world applications. * Conduct experiments to evaluate model performance and iterate on architectures and training pipelines. * Stay up-to-date with the latest research in vision-language models and incorporate advancements into projects. * Contribute to data preprocessing, augmentation, and annotation pipelines for multimodal datasets. * Document model development processes and present findings to technical and non-technical stakeholders. ## Related Videos - [Deepfakes in Realtime - How Neural Networks Are Changing Our World](https://www.wearedevelopers.com/videos/180-deepfakes-in-realtime-how-neural-networks-are-changing-our-world) - [Creating Industry ready solutions with LLM Models](https://www.wearedevelopers.com/videos/899-creating-industry-ready-solutions-with-llm-models) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Unboxing the DeepFace](https://www.wearedevelopers.com/videos/335-unboxing-the-deepface) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)