Software Engineer, RL Environments

David Joseph & Company
San Francisco, CA, United States
about 2 months ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
1 year minimum
Compensation
$180,000.0 - $220,000.0
Working hours
Regular working hours
Job source

Tech stack

Code Generation Data Structures Software Safety Software Engineering Data Pipelines

Job description

As a SWE (Environments), you’ll design the datasets and evaluation rubrics that directly influence how frontier models learn - going from hypothesis to live experiment quickly, with output feeding directly into model training runs at scale.

What you’ll be doing

  • Design data slices and explore data shapes that expose meaningful model failure modes across domains like finance, code, and enterprise workflows
  • Build and refine evaluation rubrics and reward signals for RLHF and RLVR training pipelines
  • Model annotator behavior and run experiments to improve different model capabilities
  • Develop quantitative frameworks for measuring dataset quality, diversity, and downstream impact on model alignment and capability
  • Create and manage both real-world and synthetic data pipelines
  • Partner with lab research teams to translate their training objectives into concrete data and evaluation specifications

Requirements

  • 1-4 years of software engineering experience with strong technical depth
  • Design targeted data slices that surface model failure modes across high-stakes domains (finance, code generation, enterprise workflows)
  • Build and iterate on evaluation rubrics and reward signals powering RLHF and RLVR training pipelines
  • Develop quantitative frameworks to measure dataset quality, diversity, and downstream impact on model alignment and capability
  • Own end-to-end real world and synthetic data pipelines, from scoping with research teams to production-ready evaluation specs
  • Run annotator modeling experiments to improve model capabilities across task types

Green Flags

  • Experience at RL environment companies
  • Background in AI safety or benchmarking organizations like METR or Artificial Analysis
  • Genuine obsession with how data structure, selection, and quality drive model behavior
  • Ability to design lightweight experiments and move fast
  • Former founders or early engineers at early stage startups
  • Demonstrated ability to work hard, learn fast, and care deeply about details, * Looking for standard product engineering work - the real scope is data pipelines, reward modeling, and eval infra

Benefits & conditions

  • Outsized total cash: base plus substantial profit share, plus competitive equity
  • Direct impact on frontier AI model development, working with the world’s leading AI labs
  • High ownership on a small, early team - scope, build, and ship end to end

About the company

An early-stage (post-Series A) company building the training data and evaluation infrastructure that frontier AI labs use to improve their models - designing high-signal datasets and running rigorous evaluations that go beyond static benchmarks. A small team where individual contributors have direct impact on how the next generation of models learns. The company has raised $30M (~$300M valuation), with a founding team drawn from Jane Street, Citadel, Google, Goldman, and Stanford AI Lab.

Founded 2025 ¡ 11-50 people ¡ Industry: Consumer Tech

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:08 min

Applying software engineering environments and testing to data pipelines

Matthias Niehoff Matthias Niehoff ¡ World Congress 2024

2:22 min

Reviewing and troubleshooting code generated by AI assistants

Cassidy Williams Cassidy Williams ¡ Coffee With Developers

4:42 min

Building robust data structures with structs and bound functions

Rainer Stropek Rainer Stropek ¡ World Congress 2021

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt ¡ LIVE

47 sec

Building modern data pipelines for legacy exports

Dr. Alexander Wachtel Dr. Alexander Wachtel +1 ¡ World Congress 2025

3:47 min

Understanding the trade-offs of automated code generation

Marco Podien Marco Podien ¡ World Congress 2025

Videos

See all

Related articles

See all