> Markdown version of [/jobs/ext/2188592-reinforcement-learning-engineer](https://www.wearedevelopers.com/jobs/ext/2188592-reinforcement-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Reinforcement Learning Engineer - **Company:** Bright Vision Technologies - **Location:** Shrewsbury, MA, United States (Remote available) - **Experience:** Expert - **Salary:** $100,000.0 - $150,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Neural Networks, Computer Clusters, Python (Programming Language), Machine Learning, Reinforcement Learning, Large Language Models, Multi-Agent Systems, Deep Learning, Information Technology, Free and Open-Source Software - **Published:** August 22, 2026 - **Apply:** https://www.careerjet.com/jobad/usd7371c2dcb571ca9739cf7fba5238c32 ## About the Role Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position., * Master's or PhD in Computer Science, Machine Learning, or a related field; or equivalent applied experience. * Six or more years of combined RL research and engineering experience. * Strong proficiency in Python and modern deep learning frameworks. * Hands-on experience with at least one major RL library or in-house RL stack. * Solid understanding of probability, optimization, and the theoretical foundations of RL. * Experience designing and tuning reward functions in non-trivial environments. * Familiarity with simulation environments and large-scale experience collection. * Experience training neural network policies on GPU clusters. * Strong written and verbal communication skills. * Track record of shipping or publishing impactful RL work. Preferred Qualifications * Experience with RLHF for large language models. * Familiarity with multi-agent RL or hierarchical RL. * Exposure to robotics, control systems, or autonomous driving. * Publications in RL or related research venues. * Open-source contributions to RL libraries or environments. ## Description * Design and implement reinforcement learning solutions for sequential decision-making problems in real and simulated environments. * Develop, calibrate, and maintain simulation environments suitable for large-scale agent training. * Implement and evaluate modern RL algorithms including policy gradient, actor-critic, off-policy, and offline RL methods. * Engineer reward functions and shaping strategies that align agent behavior with desired outcomes and safety constraints. * Apply offline RL and imitation learning techniques where exploration is costly or unsafe. * Use RLHF, DPO, and related techniques for fine-tuning large language models when relevant. * Build scalable training infrastructure for distributed RL, including efficient experience collection and replay systems. * Optimize training stability and sample efficiency through algorithmic and engineering improvements. * Design rigorous evaluation protocols, including out-of-distribution and adversarial test cases. * Implement safety mechanisms such as constraint enforcement, conservative policies, and human-in-the-loop oversight. * Collaborate with applied scientists and product teams to identify high-value RL use cases. * Monitor deployed policies and models in production for drift, regression, and unintended behaviors, building the alerting and dashboards that surface issues before they meaningfully affect users. * Document methodology, design decisions, and operational characteristics for internal stakeholders. * Stay current with RL research and translate promising techniques into production-ready solutions. ## Related Videos - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Getting Started with Machine Learning](https://www.wearedevelopers.com/videos/260-getting-started-with-machine-learning) - [How to develop an autonomous car end-to-end: Robotic Drive and the mobility revolution](https://www.wearedevelopers.com/videos/22-how-to-develop-an-autonomous-car-end-to-end-robotic-drive-and-the-mobility-revolution) - [Adding knowledge to open-source LLMs](https://www.wearedevelopers.com/videos/1522-adding-knowledge-to-open-source-llms) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [What non-automotive Machine Learning projects can learn from automotive Machine Learning projects](https://www.wearedevelopers.com/videos/397-what-non-automotive-machine-learning-projects-can-learn-from-automotive-machine-learning-projects) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)