> Markdown version of [/jobs/ext/731488-agent-rl-infra-engineer](https://www.wearedevelopers.com/jobs/ext/731488-agent-rl-infra-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Agent RL Infra Engineer - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $224,000.0 - **Contract:** Permanent contract - **Skills:** Training Data, Artificial Intelligence, Distributed Computing Environment, InfiniBand, Python (Programming Language), Delivery Pipeline, Machine Learning Operations, Hardware Infrastructure, Microservices - **Published:** June 29, 2026 - **Apply:** https://www.juju.com/job/00000000gcbsy5 ## About the Role + MS in CS, ML, or related field (or equivalent experience) + 10+ years of experience + Experience operationalizing fine-tuning methods (LoRA, SFT) and especially RL techniques (DPO, GRPO, PPO, RLAIF) into reusable cookbooks and self-service workflows + Familiarity with distributed training frameworks (e.g., Megatron, NeMo, DeepSpeed, FSDP, HF Accelerate) and ML ops skills covering pipeline automation, job orchestration, and GPU cluster management are important here + Proficiency in Python, Go, Rust, or similar + Background in CS, ML, or related field through formal education or equivalent experience Ways to stand out from the crowd: + Building RL environments or training recipes that other teams consumed as self-service capabilities + Familiarity with NVIDIA infrastructure (DGX, AI Factory, NVLink/InfiniBand), NeMo Microservices, or the evolving RL-for-agents ecosystem (rLLM, Agent Lightning, HUD, OpenRLHF, SkyRL) + Experience with data curation, active learning, continuous learning loops, or data flywheel architectures also valued ## Description The work splits between creating enterprise-ready RL capabilities and partnering with agent teams to put them into practice. Building RL cookbooks and environments: + Evaluate and adapt democratized RL approaches into reusable cookbooks and blueprints so agent developers can integrate self-improvement loops (GRPO, DPO, PPO, RLAIF) on their own + Design verifiable reward environments building on NeMo Gym, extending to domain-specific environments for internal use cases + Operationalize NVIDIA and third-party training backends as production services inside Sandbox + Integrate with NeMo Microservices (Curator, Customizer, Evaluator, Guardrails) to enable end-to-end data flywheel workflows for RL Infrastructure, reliability, and collaboration: + Lead data curation and active learning strategies to continuously improve training data quality + Design RL training loops for agent self-improvement: reward modeling, policy optimization, safety constraints + Integrate with AI Factory GPU infrastructure for throughput, data locality, and multi-node training + Build observability for training runs and ensure workloads meet security and governance requirements + Collaborate with platform, security, agent infrastructure, and internal customer teams on safe deployment of training outputs ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [AI That Fits Your Business, Not the Other Way Around](https://www.wearedevelopers.com/videos/100148-ai-that-fits-your-business-not-the-other-way-around) - [Microservices: how to get started with Spring Boot and Kubernetes](https://www.wearedevelopers.com/videos/242-microservices-how-to-get-started-with-spring-boot-and-kubernetes) - [Fireside Chat: Deep Learning, Deep Impact: Harnessing AI for Language Innovation](https://www.wearedevelopers.com/videos/612-fireside-chat-deep-learning-deep-impact-harnessing-ai-for-language-innovation) - [Physical AI for the Next Wave of Industrial Digitalisation](https://www.wearedevelopers.com/videos/100039-physical-ai-for-the-next-wave-of-industrial-digitalisation) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)