> Markdown version of [/jobs/ext/3086605-reinforcement-learning-engineer](https://www.wearedevelopers.com/jobs/ext/3086605-reinforcement-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Reinforcement Learning Engineer - **Company:** Bright Vision Technologies - **Location:** Monroeville, PA, United States (Remote available) - **Experience:** Expert - **Salary:** $100,000.0 - $150,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Neural Networks, Computer Clusters, Python (Programming Language), Machine Learning, Reinforcement Learning, Large Language Models, Multi-Agent Systems, Deep Learning, Information Technology, Free and Open-Source Software - **Published:** September 26, 2026 - **Apply:** https://www.careerjet.com/jobad/us3e4ec027b08722726719cde99aab392d ## About the Role Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position., * Master's or PhD in Computer Science, Machine Learning, or a related field; or equivalent applied experience. * Six or more years of combined RL research and engineering experience. * Strong proficiency in Python and modern deep learning frameworks. * Hands-on experience with at least one major RL library or in-house RL stack. * Solid understanding of probability, optimization, and the theoretical foundations of RL. * Experience designing and tuning reward functions in non-trivial environments. * Familiarity with simulation environments and large-scale experience collection. * Experience training neural network policies on GPU clusters. * Strong written and verbal communication skills. * Track record of shipping or publishing impactful RL work. Preferred Qualifications * Experience with RLHF for large language models. * Familiarity with multi-agent RL or hierarchical RL. * Exposure to robotics, control systems, or autonomous driving. * Publications in RL or related research venues. * Open-source contributions to RL libraries or environments. ## Description * Design and implement reinforcement learning solutions for sequential decision-making problems in real and simulated environments. * Develop, calibrate, and maintain simulation environments suitable for large-scale agent training. * Implement and evaluate modern RL algorithms including policy gradient, actor-critic, off-policy, and offline RL methods. * Engineer reward functions and shaping strategies that align agent behavior with desired outcomes and safety constraints. * Apply offline RL and imitation learning techniques where exploration is costly or unsafe. * Use RLHF, DPO, and related techniques for fine-tuning large language models when relevant. * Build scalable training infrastructure for distributed RL, including efficient experience collection and replay systems. * Optimize training stability and sample efficiency through algorithmic and engineering improvements. * Design rigorous evaluation protocols, including out-of-distribution and adversarial test cases. * Implement safety mechanisms such as constraint enforcement, conservative policies, and human-in-the-loop oversight. * Collaborate with applied scientists and product teams to identify high-value RL use cases. * Monitor deployed policies and models in production for drift, regression, and unintended behaviors, building the alerting and dashboards that surface issues before they meaningfully affect users. * Document methodology, design decisions, and operational characteristics for internal stakeholders. * Stay current with RL research and translate promising techniques into production-ready solutions.