Reinforcement Learning Engineer

Bright Vision Technologies
Westford, United States of America
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
$ 150K

Job location

Remote
Westford, United States of America

Tech stack

Artificial Neural Networks
Computer Clusters
Python
Machine Learning
Reinforcement Learning
Large Language Models
Multi-Agent Systems
Deep Learning
Information Technology
Free and Open-Source Software

Job description

  • Design and implement reinforcement learning solutions for sequential decision-making problems in real and simulated environments.
  • Develop, calibrate, and maintain simulation environments suitable for large-scale agent training.
  • Implement and evaluate modern RL algorithms including policy gradient, actor-critic, off-policy, and offline RL methods.
  • Engineer reward functions and shaping strategies that align agent behavior with desired outcomes and safety constraints.
  • Apply offline RL and imitation learning techniques where exploration is costly or unsafe.
  • Use RLHF, DPO, and related techniques for fine-tuning large language models when relevant.
  • Build scalable training infrastructure for distributed RL, including efficient experience collection and replay systems.
  • Optimize training stability and sample efficiency through algorithmic and engineering improvements.
  • Design rigorous evaluation protocols, including out-of-distribution and adversarial test cases.
  • Implement safety mechanisms such as constraint enforcement, conservative policies, and human-in-the-loop oversight.
  • Collaborate with applied scientists and product teams to identify high-value RL use cases.
  • Monitor deployed policies and models in production for drift, regression, and unintended behaviors, building the alerting and dashboards that surface issues before they meaningfully affect users.
  • Document methodology, design decisions, and operational characteristics for internal stakeholders.
  • Stay current with RL research and translate promising techniques into production-ready solutions.

Requirements

Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position., * Master's or PhD in Computer Science, Machine Learning, or a related field; or equivalent applied experience.

  • Six or more years of combined RL research and engineering experience.
  • Strong proficiency in Python and modern deep learning frameworks.
  • Hands-on experience with at least one major RL library or in-house RL stack.
  • Solid understanding of probability, optimization, and the theoretical foundations of RL.
  • Experience designing and tuning reward functions in non-trivial environments.
  • Familiarity with simulation environments and large-scale experience collection.
  • Experience training neural network policies on GPU clusters.
  • Strong written and verbal communication skills.
  • Track record of shipping or publishing impactful RL work.

Preferred Qualifications

  • Experience with RLHF for large language models.
  • Familiarity with multi-agent RL or hierarchical RL.
  • Exposure to robotics, control systems, or autonomous driving.
  • Publications in RL or related research venues.
  • Open-source contributions to RL libraries or environments.

About the company

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

Apply for this position