Founding Machine Learning - Eval Layer

One Inc
San Francisco, CA, United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$150,000.0 - $275,000.0
Working hours
Regular working hours
Job source

Tech stack

Information Engineering Python (Programming Language) Machine Learning Pytorch Large Language Models

Job description

  • Train evaluation models: Develop VLMs that classify and verify policy behavior.
  • Build confidence layers: Convert model outputs into trustworthy signals the customer can act on.
  • Improve model grounding: Make the eval models reason accurately about physical and spatial scenes.
  • Build a self-improving eval layer: Develop data engine that makes the eval models sharper with each customer’s deployments and corrections.

Requirements

  • Very strong coding in Python and PyTorch.
  • VLM/LLM training: Track record in training VLMs or LLMs.
  • Evals experience: Developed and shipped evals for VLMs or LLMs.

Benefits & conditions

Compensation Range: $150K - $275K

About the company

One Robot builds task-specific world models and an evaluation platform for robot manipulation policies.

Training end-to-end policies for robots is vibes-based today. Teams collect data, train, deploy on a real robot, find out what fails, collect more, retry. We replace the trial-and-error with rigorous validation that tells you where your policy will fail and what data to collect to fix it.

Robotics can’t industrialize without an evaluation layer. We’re building it.

We’re solving challenging technical problems around long-horizon autoregressive generation, world model controllability, and closing the sim-to-real gap. We work with real customer data, real failures, and real deployment pressure.

We’re based in San Francisco, backed by Accel, YC, several exited founders, and engineering leaders at leading AI companies.

We’re small and deliberately so. Everyone is an IC with deep ownership of a wide surface area. The culture is fast iteration and direct responsibility.

Hemanth Sarabu and Elton Shon co-founded One Robot after leading robot learning together at Industrial Next (YC W22), bringing experience from Google, NASA JPL, and Tesla.

We’re building the evaluation layer to understand policy failure modes before they hit production. You’ll own modeling work that makes the eval trustworthy.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

2:36 min

Applying supervised machine learning for practical rule extraction

Katja Träumner

3:32 min

Fundamentals and limitations of large language models

Krzystof Czieslak · LIVE

1:13 min

Why evaluations are the primary lever for language models

Merrill Lutsky Merrill Lutsky · World Congress 2025

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

4:42 min

Programmatic model evaluation and custom metrics via MLflow

Viktoria Semaan Viktoria Semaan · World Congress 2026 Europe

Videos

See all

Related articles

See all