> Markdown version of [/jobs/ext/173482-research-scientist-engineer-data-evaluation](https://www.wearedevelopers.com/jobs/ext/173482-research-scientist-engineer-data-evaluation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Research Scientist / Engineer - Data & Evaluation - **Company:** Rhoda ai - **Location:** Palo Alto, CA, United States - **Contract:** Permanent contract - **Skills:** Computer Vision, Data Deduplication, Language Modeling, Data Pipelines, Data Selection - **Published:** May 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=9d42f0fce4b26c8c ## About the Role Do you have experience in Scientific publications?, * Strong understanding of data-centric ML and how web video data quality affects large generative model performance * Experience building large-scale video data pipelines: ingestion, filtering, deduplication, and quality scoring * Familiarity with video-specific data characteristics: temporal structure, motion quality, scene diversity, and action content * Solid ML fundamentals with hands-on experience training or evaluating large generative models * Ability to design evaluations for video generation models that are diagnostic, reproducible, and actionable * Staff-level candidates are expected to define technical direction and drive research strategy independently; senior/MTS candidates execute complex projects with strong fundamentals and growing scope Nice to Have (But Not Required) * PhD or strong research background in ML, computer vision, or a related field * Experience with large-scale web video dataset curation (e.g., WebVid, HowTo100M, Ego4D, or similar) * Familiarity with video generation quality metrics (FVD, perceptual quality, motion consistency) * Experience running VLM or CLIP-style inference at scale for automated video filtering and annotation * Prior work on evaluation methodology for video generation or world models * Understanding of how web video data properties connect to downstream robotic action prediction * Publication record at NeurIPS, ICML, ICLR, CVPR, or related venues ## Description * Design and implement scalable curation pipelines for web-scale video pretraining data: ingestion, deduplication, quality filtering, and content classification across internet-scale video corpora * Develop video-specific annotation frameworks and quality filters - motion quality, scene diversity, action content, temporal coherence - to improve pretraining signal * Build evaluation frameworks and benchmarks to measure causal video model capabilities: prediction quality, temporal coherence, long-horizon rollout fidelity, and downstream robot task performance * Research and implement data selection, mixing, and weighting strategies that improve video generation quality and transfer to robotic control * Deploy and scale vision-language models (VLMs) and video understanding models for automated annotation, filtering, and content scoring at web scale * Collaborate closely with pre-training and post-training teams to ensure data quality and evaluation methodology drive research decisions * Track model capability trends across training runs, catching regressions and surfacing improvements early, * The video curation and evaluation rigor you build directly determines pretraining quality and research iteration speed for the entire team * Build the benchmark infrastructure that gives the team an honest signal of model progress toward real robot performance * High leverage: improvements to data quality compound across every training run * Work at the intersection of large-scale systems and generative model research with visibility across all model development ## Related Videos - [Intelligent Data Selection for Continual Learning of AI Functions](https://www.wearedevelopers.com/videos/367-intelligent-data-selection-for-continual-learning-of-ai-functions) - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [Focoos AI: Building the Future of Computer Vision](https://www.wearedevelopers.com/videos/1659-focoos-ai-building-the-future-of-computer-vision) - [Carl Lapierre - Exploring Advanced Patterns in Retrieval-Augmented Generation](https://www.wearedevelopers.com/videos/1235-carl-lapierre-exploring-advanced-patterns-in-retrieval-augmented-generation) - [Stop Guessing, Start Measuring: Evaluating RAG Systems with Synthetic Test Data](https://www.wearedevelopers.com/videos/1982-stop-guessing-start-measuring-evaluating-rag-systems-with-synthetic-test-data) - [Python-Based Data Streaming Pipelines Within Minutes](https://www.wearedevelopers.com/videos/1233-python-based-data-streaming-pipelines-within-minutes) ## Related Articles - [Dev Digest 129 - Now that's what I call private data!](https://www.wearedevelopers.com/magazine/468-dev-digest-129-now-that-s-what-i-call-private-data) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 138 - Are you secure about this?](https://www.wearedevelopers.com/magazine/486-dev-digest-138-are-you-secure-about-this) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [DeepSeek R1 vs ChatGPT o1: How Do They Compare?](https://www.wearedevelopers.com/magazine/542-deepseek-r1-vs-chatgpt-o1-how-do-they-compare) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)