> Markdown version of [/jobs/ext/1585585-senior-robot-learning-engineer](https://www.wearedevelopers.com/jobs/ext/1585585-senior-robot-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Robot Learning Engineer - **Company:** Wave Recruitment Ltd - **Location:** Bristol, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Python (Programming Language), Language Modeling, Pytorch, Transfer Learning, Stable Diffusion - **Published:** July 1, 2026 - **Apply:** https://www.gradsouthwest.com/jobs/apply/ats-redirect/?id=2850727966-2 ## About the Role + PhD/MSc in ML, Robotics, CS, or related field with 4+ years of equivalent industry research experience + Demonstrated expertise training and deploying learned manipulation policies on real robots + Strong background in at least two of: behaviour cloning, diffusion policies, VLA/VLM architectures, RL for manipulation + PyTorch and large-scale (multi-GPU, distributed) training + Track record of publications at top-tier venues (CoRL, RSS, ICRA, NeurIPS, ICML, ICLR), or equivalent demonstrated research impact through deployed systems, patents, or significant open-source contributions + Strong Python; production-quality research code with proper testing, type hints, and documentation Useful: + Hands-on experience with humanoid or bi-manual manipulation platforms + Diffusion transformer, ACT, or VLA architectures specifically + Pre-trained vision/language models for robot control (CLIP, DINOv2, PaliGemma) + MuJoCo, Isaac Sim, or ManiSkill for sim-to-real policy training + RL fine-tuning of pre-trained policies (residual RL, DPPO, or similar) ## Description This robot learning role is with a seriously exciting scale up. The platform is mature, the data is flowing, and the team is ready to scale its most promising research directions into production-grade manipulation policies. They need someone to lead the development and deployment of large behaviour models, taking diffusion transformers, VLAs, and language-conditioned policies from the literature onto a real bi-manual humanoid. This is not a research-only role. You'll inherit a mature policy training codebase, a VR teleoperation pipeline producing high-frequency multi-modal data, and a Gymnasium environment wrapping a real robot. The work you ship runs on hardware. The Role You will architect, train, and deploy end-to-end large behaviour models for bi-manual and mobile manipulation, and lead the maturing of the early-stage RL pipeline. The key responsibilities + Architect, train, and evaluate end-to-end large behaviour models for bi-manual and mobile manipulation + Advance diffusion transformer policies, mature VLA integration, and develop language conditioning for true multi-task generalisation + Apply RL to refine pre-trained policies: RL token fine-tuning, residual RL, off-policy RL with reference-action regularisation, RL-based fine-tuning of diffusion policies + Build a systematic sim-to-real transfer pipeline, connecting existing simulation infrastructure to training + Deploy and iterate learned policies on physical robot hardware + Mentor junior researchers and engineers, and publish at top-tier venues, Key contribution areas Policy Architecture & Training + End-to-end large behaviour models for bi-manual and mobile manipulation + Scale and evolve diffusion transformer policies, VLA integration, and language conditioning + Extend the imitation learning pipeline to leverage growing teleoperation datasets + Apply RL to push beyond what imitation alone can reach + Target sub-millimetre precision and contact-rich manipulation Generalisation & Scaling + Develop policies that generalise across tasks, object categories, and environments + Move from single-task to multi-task and task-conditioned architectures + Design hierarchical behaviour systems for long-horizon manipulation + Investigate data-efficient learning: few-shot adaptation, transfer learning, multi-dataset training + Drive systematic ablations across architectures Sim-to-Real & Deployment + Build the sim-to-real transfer pipeline: domain randomisation, rendering augmentation, sim-to-real benchmarking + Deploy and iterate learned policies on physical robot hardware + Extend the Gymnasium environment wrapper and integrate with the robot's control stack + Leverage perception team outputs (keypoints, learned features, 3D point clouds) for policy conditioning Research Leadership + Track the literature and bring relevant advances back to the team + Identify and propose new research directions aligned with the manipulation roadmap + Mentor junior researchers and engineers + Publish at top-tier venues - conference attendance and open-source contributions are actively supported ## Related Videos - [Robots Among us: Advances in AI for Everyday Androids](https://www.wearedevelopers.com/videos/100230-robots-among-us-advances-in-ai-for-everyday-androids) - [Your imaginations is (no longer) the limit: how Generative AI empowers people to be creative](https://www.wearedevelopers.com/videos/741-your-imaginations-is-no-longer-the-limit-how-generative-ai-empowers-people-to-be-creative) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Unveiling the Magic: Scaling Large Language Models to Serve Millions](https://www.wearedevelopers.com/videos/1619-unveiling-the-magic-scaling-large-language-models-to-serve-millions) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Robots are coming into the wild! Full-Stack Robotics Engineers, be ready!](https://www.wearedevelopers.com/videos/479-robots-are-coming-into-the-wild-full-stack-robotics-engineers-be-ready) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [A 5-Step Open-Source Setup for Agentic Engineering](https://www.wearedevelopers.com/magazine/738-a-5-step-open-source-setup-for-agentic-engineering) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)