> Markdown version of [/jobs/ext/2181895-research-engineer-computer-vision](https://www.wearedevelopers.com/jobs/ext/2181895-research-engineer-computer-vision). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Research Engineer, Computer Vision - **Company:** BINGHAMTOM UNIVERSITY - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Clean Code Principles, Artificial Intelligence, Computer Vision, Batch Processing, Cloud Computing, Metadata, Language Modeling, Object Detection, Sensor Fusion, Systems Integration, Video Editing, Web Application Frameworks, Pytorch, Deep Learning, Model Validation, Backend, Machine Learning Operations, Autodesk Autocad - **Published:** August 22, 2026 - **Apply:** https://www.themuse.com/jobs/autodesk/sr-research-engineer-computer-vision?utm_source=uconnect ## About the Role * Bachelor's degree in Computer Science, Electrical Engineering, Robotics, or related field (or equivalent practical experience) * 4+ years of experience building computer vision systems using Python * Strong experience with deep learning for computer vision (detection, segmentation, and/or video understanding) using modern frameworks such as PyTorch * Experience taking ML prototypes into reliable pipelines, including evaluation, monitoring, and failure analysis * Experience building or integrating ML systems into cloud or backend workflows (batch processing and/or services) * Strong collaboration and communication skills; ability to work across teams and stakeholders, * Experience with vision-language models (VLMs) and multimodal systems (for example: grounded vision, open-vocabulary recognition, retrieval-augmented multimodal reasoning) * Experience with multimodal fusion (combining imagery/video with metadata, documents, and sensor signals) * Experience with video pipelines (tracking, temporal aggregation, long-video processing) * Experience with real-world datasets, including data curation, labelling strategy, augmentation, and quality control under limited data constraints * Experience developing reusable platform components adopted across multiple teams What Success Looks Like * Delivered an end-to-end system that ingests real-world image/video inputs and outputs a structured, queryable set of observations (objects plus activities/events), with clear accuracy and reliability metrics * Demonstrated robustness to common visual failure modes (lighting, occlusion, clutter, camera variation) and measurable improvements when contextual signals are available * Built a modular pipeline architecture (segmentation/detection/VLM reasoning components) that can be reused and extended across domains and teams * Maintained strong engineering quality: reproducible experiments, documented decisions, maintainable code, and dependable integrations Keywords (for candidate matching) Computer Vision, Deep Learning, PyTorch, Object Detection, Segmentation, Tracking, Video Understanding, Vision-Language Models (VLM), Multimodal AI, Open-Vocabulary, Grounding, Sensor Fusion, Data Curation, Model Evaluation, Benchmarking, Cloud ML Pipelines, Batch Processing, MLOps, Observability #LI-JK3 ## Description We are hiring a Senior Software Engineer focused on Computer Vision and Multimodal AI to build robust perception and understanding systems used across multiple teams and product areas. You will develop end-to-end pipelines that transform images and video into structured, reliable observations by combining modern vision models with multimodal reasoning and contextual signals (for example: domain metadata, documents, and sensor inputs) This role blends applied research with strong software engineering: rapid iteration, rigorous evaluation, and production-minded implementation for cloud-scale batch processing and interactive workflows, * Design, build, and improve multi-stage computer vision pipelines that may include segmentation, detection, tracking, and VLM-based analysis, producing structured outputs (entities, attributes, actions/events, confidence, provenance) * Build systems that handle real-world variability in visual inputs (for example: low resolution, poor lighting, motion blur, cluttered scenes, inconsistent capture devices) * Work with diverse media types such as photos, video, timelapse, 360 video, and RGB-D when available * Fuse visual evidence with contextual inputs such as metadata, documents, and sensor streams to improve recognition quality and reduce ambiguity * Evaluate and integrate state-of-the-art vision and vision-language foundation models, including open-vocabulary recognition, grounded perception, segmentation, and multimodal reasoning * Apply fine-tuning or adaptation approaches when needed; partner with ML teams on training, data strategy, and infrastructure best practices * Define measurable acceptance criteria and benchmarking for accuracy, robustness, latency/cost, and reliability across datasets and domains * Build scalable cloud workflows for batch processing and integrate outputs with APIs and downstream consumers * Improve operational performance and cost via batching, caching, model selection, and pipeline observability * Write maintainable code, contribute to design docs, code reviews, shared libraries, and cross-team technical decisions ## Related Videos - [Machine Learning for Software Developers (and Knitters)](https://www.wearedevelopers.com/videos/154-machine-learning-for-software-developers-and-knitters) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [A Data Mesh needs Open Metadata](https://www.wearedevelopers.com/videos/505-a-data-mesh-needs-open-metadata) - [Robots 2.0: When artificial intelligence meets steel](https://www.wearedevelopers.com/videos/1452-robots-2-0-when-artificial-intelligence-meets-steel) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 129 - Now that's what I call private data!](https://www.wearedevelopers.com/magazine/468-dev-digest-129-now-that-s-what-i-call-private-data) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)