> Markdown version of [/jobs/ext/3396316-research-engineer-rl-environments-and-infrastructure](https://www.wearedevelopers.com/jobs/ext/3396316-research-engineer-rl-environments-and-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Research Engineer, RL Environments and Infrastructure - **Company:** Hyphen Hyphen LLC - **Location:** San Francisco, CA, United States - **Contract:** Permanent contract - **Skills:** Training Data, Application Programming Interfaces (APIs), Artificial Intelligence, C++ (Programming Language), Nvidia CUDA, Databases, Python (Programming Language), TypeScript, Reinforcement Learning, Rust (Programming Language), Graphics Processing Unit (GPU), Golang - **Published:** September 10, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=4de0bd5d217bf4d1 ## About the Role * Production AI Experience: Track record of deploying evaluations, RL loops, or sandboxed agent environments into production. * Systems Polyglot: Deep systems background (Python, Rust, C++, Go, TypeScript, CUDA)-you pick up new tools and frameworks within days. * San Francisco On-Site: In-person collaboration at our SF office to iterate quickly with founders and domain experts. * Pragmatic Problem Solver: Comfortable navigating raw paper implementations, undocumented SDKs, and custom distributed training setups. ## Description * Build Stateful Agent Environments: Design deterministically verifiable, stateful sandboxes (web, OS, API, database) where agents can execute 50+ step action trajectories safely. * Scale Post-Training & RL Pipelines: Implement high-throughput post-training infrastructure (RLHF, Direct Preference Optimization, Process-Supervised Reward Models) for dynamic policy optimization. * Design Enterprise Benchmarks: Formulate evaluation metrics and automated grading harnesses that catch agent drift, hallucination, and loops in realistic enterprise environments. * Systems Optimization: Keep latency low and compute efficiency high across distributed GPUs and sandboxed runtime environments. ## Related Videos - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Kubernetes and Microservices with Multi-Model Databases](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) - [Do TypeScript without TypeScript](https://www.wearedevelopers.com/videos/327-do-typescript-without-typescript) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) - [RTX AI PC: Developing local and edge AI applications](https://www.wearedevelopers.com/videos/100078-rtx-ai-pc-developing-local-and-edge-ai-applications) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Building AI Solutions with Rust and Docker](https://www.wearedevelopers.com/magazine/494-building-ai-solutions-with-rust-and-docker) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)