Software Engineer, RL Environments
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
As a SWE (Environments), youâll design the datasets and evaluation rubrics that directly influence how frontier models learn - going from hypothesis to live experiment quickly, with output feeding directly into model training runs at scale.
What youâll be doing
- Design data slices and explore data shapes that expose meaningful model failure modes across domains like finance, code, and enterprise workflows
- Build and refine evaluation rubrics and reward signals for RLHF and RLVR training pipelines
- Model annotator behavior and run experiments to improve different model capabilities
- Develop quantitative frameworks for measuring dataset quality, diversity, and downstream impact on model alignment and capability
- Create and manage both real-world and synthetic data pipelines
- Partner with lab research teams to translate their training objectives into concrete data and evaluation specifications
Requirements
- 1-4 years of software engineering experience with strong technical depth
- Design targeted data slices that surface model failure modes across high-stakes domains (finance, code generation, enterprise workflows)
- Build and iterate on evaluation rubrics and reward signals powering RLHF and RLVR training pipelines
- Develop quantitative frameworks to measure dataset quality, diversity, and downstream impact on model alignment and capability
- Own end-to-end real world and synthetic data pipelines, from scoping with research teams to production-ready evaluation specs
- Run annotator modeling experiments to improve model capabilities across task types
Green Flags
- Experience at RL environment companies
- Background in AI safety or benchmarking organizations like METR or Artificial Analysis
- Genuine obsession with how data structure, selection, and quality drive model behavior
- Ability to design lightweight experiments and move fast
- Former founders or early engineers at early stage startups
- Demonstrated ability to work hard, learn fast, and care deeply about details, * Looking for standard product engineering work - the real scope is data pipelines, reward modeling, and eval infra
Benefits & conditions
- Outsized total cash: base plus substantial profit share, plus competitive equity
- Direct impact on frontier AI model development, working with the worldâs leading AI labs
- High ownership on a small, early team - scope, build, and ship end to end
About the company
An early-stage (post-Series A) company building the training data and evaluation infrastructure that frontier AI labs use to improve their models - designing high-signal datasets and running rigorous evaluations that go beyond static benchmarks. A small team where individual contributors have direct impact on how the next generation of models learns. The company has raised $30M (~$300M valuation), with a founding team drawn from Jane Street, Citadel, Google, Goldman, and Stanford AI Lab.
Founded 2025 ¡ 11-50 people ¡ Industry: Consumer Tech
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Dev Digest 120 - Apple and peers
How We Built a Worry-Free System That Runs for 10+ Years â And What Weâd Do Again
RĂŠsumĂŠ-Driven Development: How IT trends affect the job market for software developers
Highest Paying Tech Companies for Developers