> Markdown version of [/jobs/ext/3277510-research-engineer-scientist-post-training-reinforcement-learning-london](https://www.wearedevelopers.com/jobs/ext/3277510-research-engineer-scientist-post-training-reinforcement-learning-london). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Research Engineer / Scientist, Post-training & Reinforcement Learning - London - **Company:** H Company - **Location:** London, UK - **Salary:** £56,865.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Algorithm Design, Computer Programming, Computer Literacy, Distributed Computing Environment, Distributed Systems, Python (Programming Language), Language Modeling, Node.Js, Tensorflow, Reinforcement Learning, Data Ingestion, Pytorch, Large Language Models, Deep Learning, Data Pipelines - **Published:** September 1, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5701325486 ## About the Role You have a strong research engineer / scientist mindset with experience training and improving large language models at different scales in distributed computing settings whether that's modelling, data collection, experimenting and ablating, implementing SOTA. * You have strong programming skills in Python, Rust, or similar; and strong software engineering fundamentals building performant and reliable systems. * Proficient in deep learning frameworks (Pytorch, JAX, TensorFlow). * You can work on different layers of the stack from low-level training backends, data ingestion to ML/RL algorithmic design and implementation. * You know when and where to be rigorous and slower versus when to break and iterate quickly. * You have trained LLMs/VLMs with techniques such as SFT, DPO, RLHF/RLVR, reward modelling, offline RL, distillation, etc. * You have experience with offline and online reinforcement learning in or outside of the context of language models. * Publications in top-tier AI conferences (e.g., NeurIPS, ICML, CVPR, ACL, ICCV, AAMAS, ...) * Advanced degree (PhD or MSc) in a relevant field (e.g., ML, DL, NLP, CV) * Experience with large-scale distributed training and inference (multi-node, large models, MoE, parallelism strategies, etc) * Experience training models for computer use or other multi-turn and/or multimodal agentic settings. * Extensive experience with reinforcement learning with sparse rewards. * Experience with multi-domain training, data mixture design, curriculum learning, model merging, distillation. * You are a good communicator, collaborative and low-ego. * You are able to handle a controlled-chaotic environment with a high-degree of between-teams dependencies and collaboration. * You have a go-do attitude and can balance personal conviction/interests with wider team needs. * You don't shy away from hard research or engineering problems. ## Description * Develop and train advanced LLMs and VLMs, including multimodal architectures * Research and implement training methods for enhanced capabilities like instruction following and tool use * Design and optimize data pipelines and training systems for large-scale distributed training * Collaborate with cross-functional teams to integrate models into agentic AI systems * Evaluate model performance and communicate findings to stakeholders * Stay current with advancements in LLMs, VLMs, and related fields ## Related Videos - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [Inside the Mind of an LLM](https://www.wearedevelopers.com/videos/1617-inside-the-mind-of-an-llm) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Stop Using Node.js Like It’s 2020! - Alfonso Graziano](https://www.wearedevelopers.com/videos/1863-stop-using-node-js-like-it-s-2020-alfonso-graziano) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [ I Gave a Video Editor More Autonomy Than a Trading Bot. On Purpose.](https://www.wearedevelopers.com/magazine/773-i-gave-a-video-editor-more-autonomy-than-a-trading-bot-on-purpose) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)