> Markdown version of [/jobs/ext/1499219-staff-applied-scientist-reinforcement-learning](https://www.wearedevelopers.com/jobs/ext/1499219-staff-applied-scientist-reinforcement-learning). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Applied Scientist, Reinforcement Learning - **Company:** Hippocratic AI - **Location:** Menlo Park, CA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Python (Programming Language), Node.Js, Reinforcement Learning, Pytorch, Large Language Models, Software Coding - **Published:** July 30, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=7db33aa4a9ccb2f1 ## About the Role * MS or PhD in CS or relevant field * 5+ years or experience in NLP, LLM training, or RL * 2+ years experience in RL for LLM post-training * Experience with large-scale (50B+ parameter and multi-node) LLM training * Strong Python and PyTorch coding skills * Experience with RLHF, RLVR, LLM-as-judge or similar methods for LLM post-training Nice-to-Have: * Publications at top venues (NeurIPS, ICML, ICLR, ACL, EMNLP) * Healthcare domain experience ## Description LLM post-training is where raw capability becomes reliable, safe behavior - and in healthcare, the stakes are as high as they get. You'll own the Reinforcement Learning (RL) and On-Policy Distillation (OPD) post-training pipeline end to end, to improve our models' clinical reasoning, safety, and alignment. Your models will be deployed to interact with millions of patients across diverse clinical use cases., * Design RL and OPD post-training methods (RLHF, RLVR, OPD, etc.) * Build and evaluate reward models, verifiers, and LLM-as-judge pipelines * Develop conversational AI environments and simulations for healthcare RL training with synthetic data * Automate post-training loops with agents (auto-research) * Run rigorous experiments to understand what drives post-training gains * Collaborate with research, engineering, and clinical teams ## Related Videos - [Adding knowledge to open-source LLMs](https://www.wearedevelopers.com/videos/1522-adding-knowledge-to-open-source-llms) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Unleashing the Power of Developers: Why Cybersecurity is the Missing Piece?!?](https://www.wearedevelopers.com/videos/712-unleashing-the-power-of-developers-why-cybersecurity-is-the-missing-piece) - [Building AI Applications with LangChain and Node.js](https://www.wearedevelopers.com/videos/1512-building-ai-applications-with-langchain-and-node-js) - [Inside the Mind of an LLM](https://www.wearedevelopers.com/videos/1617-inside-the-mind-of-an-llm) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)