Reinforcement Learning Engineer

Person AI Inc.
United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

Software Debugging Python (Programming Language) Machine Learning Reinforcement Learning Pytorch Deep Learning Information Technology

Job description

  • Train and iterate on reinforcement learning policies for complex grasping tasks including functional grasping, tool use, in-hand manipulation, and environment interaction.
  • Implement and refine sim-to-real transfer pipelines to bridge the gap between simulation and physical robotic hand performance.
  • Develop reward functions, curriculum strategies, and training environments in MuJoCo and Isaac Lab.
  • Run experiments on real robots alongside simulation, evaluating and debugging policy behavior on hardware.
  • Monitor, evaluate, and adapt state-of-the-art research in learning-based grasping to deploy on our humanoid platform.
  • Collaborate with the rest of the software team to deploy end-to-end grasping systems.
  • Benchmark and evaluate grasp policies across object diversity, clutter scenes, and real-world uncertainties.
  • Integrate tactile sensing and feedback into grasp policies for robust, force-aware manipulation.

Requirements

  • BS, MS, or PhD in Robotics, Computer Science, Machine Learning, or a related field.
  • 2+ years of hands-on experience in reinforcement learning for robotic manipulation; exceptional entry level candidates from relevant research labs will be considered.
  • Demonstrated ability to read, understand, and implement ideas from recent robotics and machine learning research.
  • Hands-on experience training RL agents for robotic manipulation tasks, including reward shaping and policy evaluation.
  • Experience with sim-to-real transfer: domain randomization, physics tuning, or real-world policy validation on hardware.
  • Proficiency in Python and deep learning frameworks (PyTorch, JAX), along with RL libraries such as rsl_rl or skrl.
  • Experience preparing meshes and collision geometries for RL environments in simulators such as MuJoCo and/or Isaac Sim.

Bonus Qualifications:

  • Experience deploying RL-trained policies on physical robotic hands.
  • Experience with tactile sensors and integrating tactile feedback into learned grasp policies.
  • Experience with contact-rich manipulation and force/torque estimation.
  • Familiarity with other learning-based approaches such as behavior cloning, imitation learning, or diffusion-based policy methods.
  • Publications or project work at top-tier venues (CoRL, RSS, ICRA) on grasping or dexterous manipulation.
  • Experience in a humanoid robot startup environment.

Benefits & conditions

  • We offer competitive compensation, a performance-based bonus, 99% employer covered medical benefits, early-stage equity, competitive PTO, and a company-wide paid winter break between December 24th and January 2nd.
  • You’ll shape technology that’s redefining the possibilities of robotics and human interaction.
  • Work alongside passionate teammates who value creativity, collaboration, and continuous learning.
  • Enjoy full access to advanced tools, hardware labs, and the freedom to push the boundaries of what robots can do. Persona AI is an Equal Opportunity Employer.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

1:25 min

Distinguishing artificial intelligence from deep learning

Sam Witteveen · Coffee With Developers

2:36 min

Applying supervised machine learning for practical rule extraction

Katja Träumner

1:50 min

Speaker background and introduction to applied robotics work

Carl Lapierre Carl Lapierre · WWC 2024

3:51 min

Overcoming hardware configuration barriers in machine learning

Jose Luis Latorre Millas · LIVE

4:41 min

Replacing PyTorch with ONNX runtime for AWS Lambda deployments

Marek Suppa · LIVE

Videos

See all

Related articles

See all