> Markdown version of [/jobs/ext/1920408-research-engineer-data-quality-evals](https://www.wearedevelopers.com/jobs/ext/1920408-research-engineer-data-quality-evals). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Research Engineer - Data Quality & Evals - **Company:** Epsilon, Inc. - **Location:** San Francisco, CA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Clinical Data Repository, Data Deduplication, Information Engineering, Dicom, Python (Programming Language), Language Modeling, Regression Testing, Data Processing, Large Language Models, Apache Spark, Free and Open-Source Software, Data Pipelines, Databricks - **Published:** August 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=89ac0ace7cbf7fe2 ## About the Role * 2+ years of industry or research experience in ML, data engineering, or a related area * Strong Python and solid software engineering fundamentals; comfortable building tooling and data pipelines from scratch * Strength in one or both of our core areas, with the willingness to grow into the other: + Data quality: dataset curation, filtering, deduplication, label-noise detection, or data-centric ML + Evaluation: designing metrics or eval harnesses for generative models, LLM-as-judge, or NLG / factuality evaluation * Demonstrated agency, i.e. a habit of identifying important problems and driving them to a result without waiting to be told * Comfort working in an ambiguous, fast-moving research environment and collaborating closely with research scientists Preferred Qualifications * Experience with medical imaging or clinical data (DICOM, radiology reports, clinical NLP) * Familiarity with vision-language models or multimodal training * Experience building human-in-the-loop annotation or review workflows, and reasoning about inter-annotator agreement * Experience with clinical accuracy metrics for report generation (e.g., entity / relation extraction, RadGraph-style scoring) * Experience with data pipeline and experiment tooling (Spark, Airflow, Databricks, or similar) * Publications or open-source contributions in data-centric ML, evaluation, or medical AI ## Description * Build data filtering and curation pipelines that keep VLM and classifier training sets clean, detecting label noise, misaligned image-report pairs, duplicates, corrupted studies, and low-quality samples at scale. * Develop model-based data quality signals (alignment scoring, automated flagging, active-learning loops) to surface the ambiguous or high-value cases worth human review. * Partner with radiologists and annotators to define quality criteria, adjudicate edge cases, and turn clinical judgment into reusable, scalable filters. * Design evaluation methodology for report generation that goes beyond surface-level text overlap, measuring clinical accuracy through entity and relation extraction, hallucination and omission rates, and adherence to reporting style. * Build and maintain clinical benchmark sets, stratified by modality, pathology, and difficulty, with rigorous attention to train/eval contamination. * Develop and validate model-based evaluators (LLM-as-judge, rubric grading) against radiologist judgment, and track how offline eval correlates with production and clinical outcomes. * Build continuous evaluation and regression testing so the team can measure every model change quickly and trust the result. * Work across the research stack (data, training, and evaluation) finding bottlenecks and shipping the tooling that lets research scientists move faster. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) - [ZEISS & Microsoft - Building the Next Generation Medical Ecosystem in the Cloud](https://www.wearedevelopers.com/videos/424-zeiss-microsoft-building-the-next-generation-medical-ecosystem-in-the-cloud) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [How E.On productionizes its AI model & Implementation of Secure Generative AI.](https://www.wearedevelopers.com/videos/623-how-e-on-productionizes-its-ai-model-implementation-of-secure-generative-ai) ## Related Articles - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)