> Markdown version of [/jobs/ext/2259359-rl-environment-data-engineer-researcher](https://www.wearedevelopers.com/jobs/ext/2259359-rl-environment-data-engineer-researcher). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # RL Environment Data Engineer / Researcher - **Company:** Eigent AI - **Location:** Greater London, UK - **Contract:** Permanent contract - **Skills:** Training Data, Artificial Intelligence, Code Generation, Data Auditing, Data Cleansing, Software Debugging, Python (Programming Language), Reinforcement Learning, Large Language Models, Software Coding, Code Restructuring, Data Pipelines - **Published:** August 26, 2026 - **Apply:** https://www.collegerecruiter.com/job/2815241792-rl-environment-data-engineer--researcher ## About the Role * - Strong coding skills, especially in Python, with the ability to independently build data pipelines, environments, and evaluation tools. * - Proficiency with AI coding tools for code generation, debugging, refactoring, and rapid experimentation. * - Solid understanding of reinforcement learning, post-training, reward function design, environment design, and data evaluation. * - Ability to translate real-world tasks into trainable and measurable RL environments. * Experience with data scraping, data cleaning, annotation, or data quality assessment is preferred. * - Experience with LLM agents, RLHF/RLAIF, coding agents, automated evaluation, or benchmark construction is a strong plus. * - Strong experimental mindset and engineering execution, with the ability to continuously improve systems based on data and evaluation results. ## Description We are looking for an RL Environment Data Engineer / Researcher to design, build, and refine reinforcement learning training environments across different domains. This role will focus on data collection, task definition, reward design, evaluation criteria, anti-reward-hacking mechanisms, and post-training validation of environment data effectiveness. Responsibilities * - Design and improve RL training environments across various task domains. * - Collect, clean, structure, and evaluate data used for RL environment construction and model post-training. * - Define task objectives, reward functions, and evaluation standards to ensure reliable and reproducible training signals. * - Develop technical approaches to prevent reward hacking and identify loopholes in reward design. * - Build validation environments to assess the effectiveness of post-training data and RL environment design. * - Collaborate with research, engineering, and data teams to improve environment coverage, task difficulty, and evaluation reliability. * - Follow research progress in RL environments, data evaluation, AI agents, and post-training methods, and apply relevant findings to production workflows. Requirements * - Strong coding skills, especially in Python, with the ability to independently build data pipelines, environments, and evaluation tools. * - Proficiency with AI coding tools for code generation, debugging, refactoring, and rapid experimentation. * - Solid understanding of reinforcement learning, post-training, reward function design, environment design, and data evaluation. * - Ability to translate real-world tasks into trainable and measurable RL environments. * Experience with data scraping, data cleaning, annotation, or data quality assessment is preferred. * - Experience with LLM agents, RLHF/RLAIF, coding agents, automated evaluation, or benchmark construction is a strong plus. * - Strong experimental mindset and engineering execution, with the ability to continuously improve systems based on data and evaluation results. ## Related Videos - [A walkthrough on Responsible AI Frameworks and Case Studies](https://www.wearedevelopers.com/videos/509-a-walkthrough-on-responsible-ai-frameworks-and-case-studies) - [How E.On productionizes its AI model & Implementation of Secure Generative AI.](https://www.wearedevelopers.com/videos/623-how-e-on-productionizes-its-ai-model-implementation-of-secure-generative-ai) - [RPA in the Public Sector](https://www.wearedevelopers.com/videos/86-rpa-in-the-public-sector) - [Fireside Chat: Deep Learning, Deep Impact: Harnessing AI for Language Innovation](https://www.wearedevelopers.com/videos/612-fireside-chat-deep-learning-deep-impact-harnessing-ai-for-language-innovation) - [Algorithmic Bias- Preventing Unfairness in your Algorithms](https://www.wearedevelopers.com/videos/43-algorithmic-bias-preventing-unfairness-in-your-algorithms) - [Exploring 5 Key Applications of AI Abundance with Blockchain Assurance](https://www.wearedevelopers.com/videos/971-exploring-5-key-applications-of-ai-abundance-with-blockchain-assurance) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)