RL Environment Data Engineer / Researcher

Eigent AI
Greater London, UK
2 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Training Data Artificial Intelligence Code Generation Data Auditing Data Cleansing Software Debugging Python (Programming Language) Reinforcement Learning Large Language Models Software Coding Code Restructuring Data Pipelines

Job description

We are looking for an RL Environment Data Engineer / Researcher to design, build, and refine reinforcement learning training environments across different domains. This role will focus on data collection, task definition, reward design, evaluation criteria, anti-reward-hacking mechanisms, and post-training validation of environment data effectiveness.

Responsibilities

    • Design and improve RL training environments across various task domains.
    • Collect, clean, structure, and evaluate data used for RL environment construction and model post-training.
    • Define task objectives, reward functions, and evaluation standards to ensure reliable and reproducible training signals.
    • Develop technical approaches to prevent reward hacking and identify loopholes in reward design.
    • Build validation environments to assess the effectiveness of post-training data and RL environment design.
    • Collaborate with research, engineering, and data teams to improve environment coverage, task difficulty, and evaluation reliability.
    • Follow research progress in RL environments, data evaluation, AI agents, and post-training methods, and apply relevant findings to production workflows.

Requirements

    • Strong coding skills, especially in Python, with the ability to independently build data pipelines, environments, and evaluation tools.
    • Proficiency with AI coding tools for code generation, debugging, refactoring, and rapid experimentation.
    • Solid understanding of reinforcement learning, post-training, reward function design, environment design, and data evaluation.
    • Ability to translate real-world tasks into trainable and measurable RL environments.
  • Experience with data scraping, data cleaning, annotation, or data quality assessment is preferred.
    • Experience with LLM agents, RLHF/RLAIF, coding agents, automated evaluation, or benchmark construction is a strong plus.
    • Strong experimental mindset and engineering execution, with the ability to continuously improve systems based on data and evaluation results.

Requirements

    • Strong coding skills, especially in Python, with the ability to independently build data pipelines, environments, and evaluation tools.
    • Proficiency with AI coding tools for code generation, debugging, refactoring, and rapid experimentation.
    • Solid understanding of reinforcement learning, post-training, reward function design, environment design, and data evaluation.
    • Ability to translate real-world tasks into trainable and measurable RL environments.
  • Experience with data scraping, data cleaning, annotation, or data quality assessment is preferred.
    • Experience with LLM agents, RLHF/RLAIF, coding agents, automated evaluation, or benchmark construction is a strong plus.
    • Strong experimental mindset and engineering execution, with the ability to continuously improve systems based on data and evaluation results.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:59 min

Evolving roles in AI driven software teams

Ignacio Riesgo Ignacio Riesgo +1 · World Congress 2024

2:17 min

Balancing model training with data preparation realities

Lukas Kölbl · LIVE

3:16 min

Advantages of reproducible configurations and instantaneous rollbacks

Álvaro Martín Lozano · LIVE

2:20 min

Architecting language translation with focused training data

Jaroslaw Kutylowski Jaroslaw Kutylowski +1 · World Congress 2023

3:24 min

Harmonizing federal data records with automated software robots

Clemens Schwaiger · LIVE

3:08 min

Conducting data audits and adversarial testing on models

Toju Duke · World Congress 2022

Videos

See all

Related articles

See all