> Markdown version of [/jobs/ext/2157572-research-engineer-reinforcement-learning](https://www.wearedevelopers.com/jobs/ext/2157572-research-engineer-reinforcement-learning). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Research Engineer (Reinforcement Learning) - **Company:** Jobgether - **Location:** Netherlands - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Information Leak Prevention, Programming Tools, Python (Programming Language), Machine Learning, Raw Data, Software Deployment, Reinforcement Learning, Graphics Processing Unit (GPU), Data Pipelines, Data Generation - **Published:** August 21, 2026 - **Apply:** https://www.adzuna.nl/details/5851198088 ## About the Role * Strong Python engineering skills and the ability to build reliable, production-quality systems. * Demonstrated experience taking a machine learning model from raw data through experimentation and into production. * A strong data-centric mindset, with attention to coverage, diversity, quality, and data leakage. * The ability to anticipate reward exploitation and design robust rewards, verifiers, and evaluation mechanisms. * Practical experience working with GPUs and a realistic understanding of their capabilities and limitations. * Strong judgment around when model training is the right solution-and when a simpler approach is preferable. * Ability to collaborate effectively within a remote, distributed, and highly autonomous team. * Experience with post-training techniques such as fine-tuning, reward design, or reinforcement learning, including approaches such as GRPO, is highly desirable. * Familiarity with RL and fine-tuning frameworks such as TRL, verl, OpenRLHF, or custom training loops is a plus. * Experience with technologies such as vLLM or SGLang for fast rollouts and FSDP for multi-GPU training is advantageous. * Experience training tool-using or multi-turn agents, as well as building execution sandboxes, verifiers, evaluation harnesses, or developer tooling, is valuable. * Familiarity with open-weight model families such as Qwen or Llama and techniques such as LoRA is a plus. ## Description This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Research Engineer (Reinforcement Learning) based in Netherlands. Join a small, senior engineering team building the next generation of voice- and text-driven AI agents. You'll focus on post-training models to make agents more capable, reliable, and effective over long-running interactions. Your work will span environments, verifiers, synthetic data, training experiments, evaluations, and production deployment. You'll tackle challenging problems such as persistent context, reliable tool use, and multi-turn agent behavior. The role combines hands-on research and engineering, with a strong emphasis on measurable improvements in model performance. You'll work closely with experienced engineers in a remote, collaborative environment where technical craft and creativity are highly valued. Your contributions will directly shape AI systems operating at significant production scale. Accountabilities * Build training environments, verifiers, and supporting infrastructure for post-training models. * Own the synthetic data pipeline from data generation through quality assurance and validation. * Run end-to-end training experiments, analyze results, and clearly identify the factors driving model improvements. * Design and maintain evaluations that models must pass before production releases. * Select and adapt suitable open-weight foundation models for specific agent and product requirements. * Develop trained behaviors that perform consistently across both voice and text-based agents. * Deploy trained models to production and continuously improve them based on real-world usage and feedback. * Develop robust approaches to long-horizon interactions, accumulated context, and reliable tool use during live conversations. ## Related Videos - [Data Governance in the Era of AI](https://www.wearedevelopers.com/videos/1622-data-governance-in-the-era-of-ai) - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [Enhancing AI-based Robotics with Simulation Workflows](https://www.wearedevelopers.com/videos/472-enhancing-ai-based-robotics-with-simulation-workflows) - [Adding knowledge to open-source LLMs](https://www.wearedevelopers.com/videos/1522-adding-knowledge-to-open-source-llms) - [Why Your AI Agent Keeps Hallucinating Your Data: Building Deterministic Context Layers](https://www.wearedevelopers.com/videos/2055-why-your-ai-agent-keeps-hallucinating-your-data-building-deterministic-context-layers) - [Bringing Clarity to Event Streams: Enabling Analytics and AI Through Rich Metadata](https://www.wearedevelopers.com/videos/1616-bringing-clarity-to-event-streams-enabling-analytics-and-ai-through-rich-metadata) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [The 13 Best Python Libraries for Developers in 2025](https://www.wearedevelopers.com/magazine/371-the-13-best-python-libraries-for-developers-in-2025) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)