> Markdown version of [/jobs/ext/1956563-machine-learning-researcher](https://www.wearedevelopers.com/jobs/ext/1956563-machine-learning-researcher). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Researcher - **Company:** Datalab Inc - **Location:** United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Training Data, Artificial Intelligence, Data Analysis, Data Files, Systems Analysis, Machine Learning, Open Source Technology, Systems Development Life Cycle, System Testing, Reinforcement Learning, Multi-Agent Systems, Model Validation, Build Management, Information Technology, Free and Open-Source Software, Machine Learning Operations, Data Generation - **Published:** August 6, 2026 - **Apply:** https://www.careerbuilder.com/job-details/machine-learning-researcher-rl-agentic-systems--d4ff80a3-11f8-4434-8ed5-91eb847e1927 ## About the Role * PhD or equivalent Master's Degree + 4+ years industry experience in machine learning, computer science, statistics, engineering, mathematics, economics, or related quantitative fields. * Strong understanding of AI model training pipelines, evaluation methodology, and the role of data in shaping model performance. * Experience working with large, unstructured, or semi-structured datasets used to train or evaluate ML systems. * Experience with reinforcement learning, sequential decision-making, agentic systems, tool-using models, or multi-step model evaluation. * Experience designing tasks, benchmarks, environments, simulations, or evaluation frameworks for real-world model behavior. * Strong intuition for realism, coverage, difficulty, fidelity, and meaningful outcome structure in datasets. * Strong experimental design, evaluation, benchmarking, and data-validation skills. * High ownership and ability to independently identify and solve high-impact problems. Nice to have * Experience developing evaluation frameworks or performance metrics for datasets, agentic systems, or training data. * Experience translating real-world workflows into structured tasks or environments for model evaluation. * Experience with RLHF, RLAIF, imitation learning, reward modeling, online or offline RL, or related methods. * Experience with Harbor or other agent evaluation frameworks. * Publications or open-source contributions in reinforcement learning, agents, evaluation, or data-centric AI. * Experience collaborating cross-functionally with product, infrastructure, or partnership teams. * Experience with synthetic data generation, trajectory generation, or simulation-based environments., Analysis Skills, Artificial Intelligence (AI), Benchmarking, Best Practices, Computer Science, Concrete, Construction, Continuous Improvement, Cross-Functional, Data Analysis, Data Modeling, Data Quality, Data Sets, Diversity, Economics, Experiment Design, Machine Learning, Machine Tool, Mathematics, Open Source, Performance Metrics, Performance Modeling, Problem Solving Skills, Production Systems, Publications, Reinforcement Learning, Research Skills, Scalable System Development, Scorecarding, Statistics, Systems Analysis, Training Data Sets, Training/Teaching, Workflow Analysis ## Description We're seeking a Machine Learning Researcher focused on RL and agentic systems to help define, design, and evaluate the datasets, tasks, environments, and benchmarks used to assess advanced AI systems. In this role, you'll work closely with research and engineering teams to translate real-world workflows into high-value datasets and evaluation assets: structured tasks, interactive environments, benchmark suites, and quality scorecards that help us understand how models perform in realistic settings. You'll help define what "high-quality agentic data" means in practice, using statistical, computational, and ML-driven methods to evaluate dataset quality, task design, environment fidelity, and downstream model performance. You'll work on the core problems of benchmarking real-world data, measuring how well models perform on that data, and designing RL-style or agentic environments that capture the structure of meaningful work. This is an ideal role for someone with a strong machine learning background who is excited by reinforcement learning, agentic systems, evaluation, and the role of data in shaping model behavior. You should be excited by the opportunity to build the datasets and benchmarks that help define what high-quality real-world data looks like for frontier AI systems. What You'll Do Design and build datasets, tasks, and environments Design and build datasets, tasks, environments, and evaluation assets for benchmarking agentic systems and multi-step model behavior. Translate real-world workflows into structured tasks, interaction traces, trajectories, stateful environments, and verifiable outcomes that can be used to evaluate advanced AI systems. Develop frameworks for evaluating real-world data quality Develop frameworks that assess diversity, realism, coverage, fidelity, informativeness, and downstream usefulness of datasets for agentic systems. Build quality scorecards and evaluation methods that make dataset strengths, weaknesses, and failure modes legible across teams. Benchmark model behavior in RL and agentic settings Evaluate planning, tool use, robustness, recovery from failure, task completion, and generalization behavior in RL-style or agentic environments. Connect model failures back to concrete dataset, environment, or task-design gaps and recommend improvements grounded in empirical evidence. Build scalable evaluation and validation tooling Contribute to tools and systems that automate dataset validation, environment generation, rollout analysis, benchmark construction, and evaluation workflows. Improve internal infrastructure for reproducible experimentation, benchmark management, and evaluation quality. Partner across research, engineering, and product Collaborate closely with research and engineering teams to identify data bottlenecks, improve evaluation methodology, and shape internal best practices around task-grounded AI training data. Represent DataLab's perspective in cross-functional discussions around dataset quality, benchmark design, and frontier agentic-system evaluation. What Success Looks Like Near-term: establish a strong evaluation baseline Create clear benchmark frameworks, evaluation assets, and dataset-quality scorecards that help Protege reason about how real-world data impacts advanced agentic systems. Use rigorous evaluation methods to identify meaningful dataset improvements, improve benchmark fidelity, and sharpen the company's understanding of what high-impact agentic data actually looks like in practice. ## Related Videos - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Machine learning 101: Where to begin?](https://www.wearedevelopers.com/videos/1014-machine-learning-101-where-to-begin) - [Introduction to TXT](https://www.wearedevelopers.com/videos/30-introduction-to-txt) - [Fireside Chat: Deep Learning, Deep Impact: Harnessing AI for Language Innovation](https://www.wearedevelopers.com/videos/612-fireside-chat-deep-learning-deep-impact-harnessing-ai-for-language-innovation) - [Implementing continuous delivery in a data processing pipeline](https://www.wearedevelopers.com/videos/73-implementing-continuous-delivery-in-a-data-processing-pipeline) - [Exploring 5 Key Applications of AI Abundance with Blockchain Assurance](https://www.wearedevelopers.com/videos/971-exploring-5-key-applications-of-ai-abundance-with-blockchain-assurance) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)