> Markdown version of [/jobs/ext/617185-ai-researcher-reinforcement-learning](https://www.wearedevelopers.com/jobs/ext/617185-ai-researcher-reinforcement-learning). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Researcher - Reinforcement Learning - **Company:** 1X, LLC - **Location:** San Carlos, CA, United States - **Salary:** $200,000.0 - $300,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Information Engineering, Python (Programming Language), Reinforcement Learning, Pytorch, Build Tools, Codebase - **Published:** June 18, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=1a06e785d93fba81 ## About the Role Do you have experience in Simulated training environment?, * Strong Python and/or C++ with experience in large codebases and build tools (Bazel or equivalent) * Proficiency with PyTorch for RL policy training and experimentation * Hands-on experience with simulation platforms (Isaac Sim, MuJoCo, or equivalent) for policy training at scale * Demonstrated experience training RL policies for manipulation or locomotion tasks, including addressing the sim-to-real gap on physical hardware * Preferred Skills * Experience with model-based RL or world-model-guided policy learning that leverages predictive models to improve sample efficiency * Familiarity with imitation learning or learning from demonstration (behavior cloning, GAIL, IQL) as a complement or bootstrap to RL * Experience deploying RL-trained policies to physical robots in production environments, including monitoring, failure analysis, and iterative improvement * Background in legged locomotion, dexterous manipulation, or contact-rich control for physical systems ## Description * Train and deploy RL policies for manipulation and locomotion tasks that perform reliably in real-world home environments measured by field task success rates, not just simulation benchmarks * Advance sim-to-real transfer techniques that measurably narrow the gap between simulation training performance and real-world policy behavior, enabling faster iteration cycles * Build training and evaluation infrastructure that lets the team iterate on policies faster with standardized benchmarks, automated regression detection, and clear connections between training metrics and field performance * Partner with hardware, controls, data, and QA teams to ship RL-trained skills to production customer sites, owning the handoff from research to deployment Key Competencies * Sim-to-real practitioner closing the sim-to-real gap on physical systems; understands domain randomization, reward shaping, and the engineering required to make simulated policies transfer reliably to real hardware * RL algorithms depth with strong foundation in RL algorithms (PPO, SAC, TD-MPC, or similar); can choose the right approach for the task and modify or extend it when standard methods fall short * Full-stack ownership owning data engineering, model architecture, and deployment; treats a promising training curve as the beginning of the job, not the end * Effective cross-functional partner working closely with hardware, controls, QA, and data teams to translate RL research into deployed robot skills, and communicates technical constraints clearly across disciplines ## Related Videos - [How Robots Learn to be Robots](https://www.wearedevelopers.com/videos/1632-how-robots-learn-to-be-robots) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Getting to Know Your Legacy (System) with AI-Driven Software Archeology](https://www.wearedevelopers.com/videos/1437-getting-to-know-your-legacy-system-with-ai-driven-software-archeology) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Robots Among us: Advances in AI for Everyday Androids](https://www.wearedevelopers.com/videos/100230-robots-among-us-advances-in-ai-for-everyday-androids) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [How to start an AI project for a good cause and boost your career](https://www.wearedevelopers.com/magazine/15-how-to-start-an-ai-project-for-a-good-cause-and-boost-your-career) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [DeepSeek R1 vs ChatGPT o1: How Do They Compare?](https://www.wearedevelopers.com/magazine/542-deepseek-r1-vs-chatgpt-o1-how-do-they-compare) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)