> Markdown version of [/jobs/ext/2716459-machine-learning-engineer-reinforcement-learning](https://www.wearedevelopers.com/jobs/ext/2716459-machine-learning-engineer-reinforcement-learning). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer - Reinforcement Learning - **Company:** Pony.AI, Inc. - **Location:** Fremont, United States - **Salary:** $150,000.0 - $250,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Big Data, Distributed Computing Environment, Python (Programming Language), Machine Learning, Pytorch, Large Language Models, Deep Learning, Information Technology, Machine Learning Operations - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/machine-learning-engineer-reinforcement-learning-ponyai-8000933 ## About the Role * M.S. or Ph.D. in Computer Science, Machine Learning, AI, or a related field-or equivalent practical experience. * Hands-on experience building and applying ML in production-grade settings, with a strong RL component (policy learning, preference/feedback optimization, or offline/online RL pipelines). * Depth in deep learning, sequence modeling, and generative models. * Demonstrated impact via strong publications or a clear history of shipping impactful ML systems end-to-end. * Experience with large-scale distributed training and large-scale data processing. * Ability to lead ambiguous technical work from problem framing through reliable delivery. Preferred * Background in autonomous vehicles, robotics, or complex simulation environments. * Strong grasp of modern RL and post-training techniques in LLM, dLLM, VLA and video generations. * Hands-on integration of simulation platforms with ML training and evaluation workflows. * Python fluency and frameworks such as PyTorch * Experience defining and operating metrics for complex, safety-critical AI systems. * Technical leadership: influencing stakeholders, aligning teams, and raising the bar for evaluation rigor. * Excellent communication-simple explanations of complex trade-offs. ## Description * Build scalable systems for training and fine-tuning large generative models that produce realistic, informative driving behaviors for evaluation and scenario coverage. * Implement and iterate on RL-style methods: algorithms, reward / preference objectives, and training setups suited to high-fidelity, insightful behaviors in simulation-aligned workflows (closed-loop evaluation mindset). * Ship deep learning solutions (including LLM / VLM where appropriate) that improve human-led triaging, automate high-volume workflows, and support nuanced analysis of self-driving behavior to surface critical anomalies. * Own production-oriented ML for fleet-scale assessment: training, optimization, monitoring, and iteration of models used to judge performance across large real-world exposure. * Design and evolve data + evaluation systems inspired by RL from human preferences (RLHF) and related paradigms-turning preference/judgment signals into repeatable, scalable training and evaluation loops. * Partner broadly with teams such as Prediction, Planning, Research, and platform/engineering leads to land cross-cutting improvements with clear metrics. ## Related Videos - [Getting Started with Machine Learning](https://www.wearedevelopers.com/videos/260-getting-started-with-machine-learning) - [Adding knowledge to open-source LLMs](https://www.wearedevelopers.com/videos/1522-adding-knowledge-to-open-source-llms) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How machine learning can help us tell fact from fiction](https://www.wearedevelopers.com/magazine/509-how-machine-learning-can-help-us-tell-fact-from-fiction) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)