Research Engineer, QC Automation
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
HUD builds infrastructure for companies creating training data for AI agents. As demand grows, we need robust systems that can maintain and scale data quality. As a Research Engineer, QC Automation , you’ll build the systems that make that possible. You’ll develop automated quality-control infrastructure grounded in human judgment and a deep understanding of what makes training data useful-not simply by relying on LLMs to judge other LLMs. This is a highly autonomous role for someone who enjoys ambiguous problems, learns quickly, and can turn unclear quality requirements into measurable, scalable systems. What You’ll Do
- Build quality-control systems grounded in human judgment and a deep understanding of data quality, with limited reliance on LLM-based evaluation.
- Define, formalize, and enforce quality standards for training data.
- Design experiments, benchmarks, and metrics to evaluate agent performance.
- Partner with data vendors to identify agent failure modes , debug quality issues, and improve data-generation processes.
- Build systems for auditing supplier datasets , including sampling strategies, rule-based validation, model-assisted validation, and feedback loops.
- Integrate QC insights into HUD’s infrastructure and data-vendor portal to reduce anomalies, inconsistencies, and edge cases.
- Work across unfamiliar domains and turn loosely defined quality problems into reliable, repeatable processes.
Requirements
Experience: Technical aptitude and learning potential matter more than years of experience., + Strong proficiency in Python, Docker, and Linux environments.
- Excellent judgment about what constitutes good data and how to measure it.
- Genuine curiosity about unfamiliar domains, with the ability to ask the right questions to develop a deep understanding quickly.
- Experience building scalable data-validation pipelines, automated QA/QC systems, or similar infrastructure without a prescribed roadmap.
- Experience designing or working with benchmarks and evaluations.
- Comfort working in an early-stage startup environment , taking ownership and executing independently.
-
A track record of learning quickly and tackling problems that don#39;t come with clear instructions. Strong Signals
- Knowledge of statistics and experimental design.
- Strong written and verbal communication skills.
- Comfort designing metrics, experiments, and QA/QC processes.
- Ability to construct tasks for new evaluations and benchmarks.
- Comfort operating in unstructured problem spaces.
- A tendency to dig into the underlying problem rather than reaching for the first available tool or solution.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Résumé-Driven Development: How IT trends affect the job market for software developers
Dev Digest 121 - AI goes offline
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
Building AI Solutions with Rust and Docker