Research Engineer (Reinforcement Learning)

Jobgether
Netherlands
6 days ago
Apply on www.adzuna.nl
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Information Leak Prevention Programming Tools Python (Programming Language) Machine Learning Raw Data Software Deployment Reinforcement Learning Graphics Processing Unit (GPU) Data Pipelines Data Generation

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Research Engineer (Reinforcement Learning) based in Netherlands.

Join a small, senior engineering team building the next generation of voice- and text-driven AI agents. You’ll focus on post-training models to make agents more capable, reliable, and effective over long-running interactions. Your work will span environments, verifiers, synthetic data, training experiments, evaluations, and production deployment. You’ll tackle challenging problems such as persistent context, reliable tool use, and multi-turn agent behavior. The role combines hands-on research and engineering, with a strong emphasis on measurable improvements in model performance. You’ll work closely with experienced engineers in a remote, collaborative environment where technical craft and creativity are highly valued. Your contributions will directly shape AI systems operating at significant production scale. Accountabilities

  • Build training environments, verifiers, and supporting infrastructure for post-training models.
  • Own the synthetic data pipeline from data generation through quality assurance and validation.
  • Run end-to-end training experiments, analyze results, and clearly identify the factors driving model improvements.
  • Design and maintain evaluations that models must pass before production releases.
  • Select and adapt suitable open-weight foundation models for specific agent and product requirements.
  • Develop trained behaviors that perform consistently across both voice and text-based agents.
  • Deploy trained models to production and continuously improve them based on real-world usage and feedback.
  • Develop robust approaches to long-horizon interactions, accumulated context, and reliable tool use during live conversations.

Requirements

  • Strong Python engineering skills and the ability to build reliable, production-quality systems.
  • Demonstrated experience taking a machine learning model from raw data through experimentation and into production.
  • A strong data-centric mindset, with attention to coverage, diversity, quality, and data leakage.
  • The ability to anticipate reward exploitation and design robust rewards, verifiers, and evaluation mechanisms.
  • Practical experience working with GPUs and a realistic understanding of their capabilities and limitations.
  • Strong judgment around when model training is the right solution-and when a simpler approach is preferable.
  • Ability to collaborate effectively within a remote, distributed, and highly autonomous team.
  • Experience with post-training techniques such as fine-tuning, reward design, or reinforcement learning, including approaches such as GRPO, is highly desirable.
  • Familiarity with RL and fine-tuning frameworks such as TRL, verl, OpenRLHF, or custom training loops is a plus.
  • Experience with technologies such as vLLM or SGLang for fast rollouts and FSDP for multi-GPU training is advantageous.
  • Experience training tool-using or multi-turn agents, as well as building execution sandboxes, verifiers, evaluation harnesses, or developer tooling, is valuable.
  • Familiarity with open-weight model families such as Qwen or Llama and techniques such as LoRA is a plus.

Benefits & conditions

  • Opportunity to make a significant impact on a fast-growing developer platform and help shape its future.
  • Collaboration with a small, highly experienced team that values technical excellence, creativity, and ownership.
  • Competitive salary and equity package.
  • Health, dental, and vision benefits.
  • Flexible vacation policy.
  • Remote-friendly working environment with flexibility and autonomy.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.nl
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:09 min

Challenges of interpreting raw data with language models

Clemens Vasters Clemens Vasters · World Congress 2025

6:08 min

Applying software engineering environments and testing to data pipelines

Matthias Niehoff Matthias Niehoff · World Congress 2024

1:46 min

Overcoming data scarcity with synthetic data generation

Anshul Jindal Anshul Jindal +1 · World Congress 2026 Europe

1:02 min

Training models with reward functions and reinforcement learning

Carl Lapierre Carl Lapierre · World Congress 2024

3:42 min

Building data pipelines and managing machine learning features

Hauke Brammer · World Congress 2021

2:09 min

Uncovering hidden coordinate manipulation communities in binary data

Nolan Royalty · Coffee With Developers

Videos

See all

Related articles

See all