Forward Deployed Machine Learning Engineer

Ai, Inc
United States
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Data Analysis Large Language Models Backend Data Pipelines

Job description

We’re hiring a Forward Deployed Machine Learning Engineer in our Benchmarks and Evaluations vertical. You’ll be the first MLE dedicated to this vertical and will work directly with the GM and our researchers to scale Protege’s position as a renowned leader in the space.

At Protege, we believe that real world data is one of the largest bottlenecks to AI progress. Our data and data expertise position us to be neutral arbiters for the market, helping model builders understand the current performance of their models, identify what data will improve performance, and show that improvement over time. Benchmarks and evaluations power that cycle. As an early engineer in the Benchmarks and Evaluations vertical, this role is an opportunity to help build the technical foundation for a critical area that greatly benefits current and future customers.

Requirements

  • 4+ years of engineering experience

  • Hands-on ML work evaluating models

  • Have previously owned backend and infrastructure

  • High ambiguity tolerance and bias to action

  • Comfort working with urgency to meet the pace and volume of the market demands

  • Strong written communication

Nice to Haves

  • Prior experience building benchmarks, evals, or human data pipelines for LLMs

  • Time at a frontier lab, an eval-focused team, or a research org

  • Founding or early engineer experience at a fast-moving startup

  • Familiarity with agentic systems, RL environments, code-execution sandboxes, TEE/TREs

Benefits & conditions

  • Partner with the GM and early customers to define what constitutes strong evals in different domains

  • Work with Protege researchers to design and build benchmarks

  • Build the standards on how different modalities should be processed

Own infrastructure

  • Build the backend the vertical runs on which includes data pipelines, execution environments, storage, and orchestration

  • Stand up sandboxed environments for agentic evals, where models need tools, code execution, or multi-step tasks

Go from fast iteration to product

  • Find repeatable eval patterns, infrastructure gaps, and product opportunities from live engagements

  • Partner with DataLab (our research team) on domain-specific data and research questions

What Success Looks Like

In the first 90 days, we expect the following:

  • Build an understanding of the evals landscape, the GM’s strategy, and customer demand

  • Build an understanding of what our platform and data partners can support today, and where the gap is for eval building

  • Identify the largest technical bets and ship multiple iterations of the eval infrastructure

  • Own the engineering portion of customer engagements end to end

About the company

We are building Protege to solve the biggest unmet need in AI - getting access to the right training data. The process today is time intensive, incredibly expensive, and often ends in failure. The Protege platform facilitates the secure, efficient, and privacy-centric exchange of AI training data.

Solving AI’s data problem is a generational opportunity. We’re backed by world-class investors and already powering partnerships with some of the most ambitious teams in AI. The company that succeeds will be one of the largest in AI - and in tech.

We’re a lean, fast-moving, high-trust team of builders who are obsessed with velocity and impact. Our culture is built for people who thrive on ambiguity, own outcomes, and want to shape the future of data and AI.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

1:48 min

Automating exploratory data analysis within training pipelines

Dora Petrella · World Congress 2023

6:08 min

Applying software engineering environments and testing to data pipelines

Matthias Niehoff Matthias Niehoff · World Congress 2024

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:12 min

Choosing TypeScript for complex backend applications

Maximilian Otto Maximilian Otto · World Congress 2024

Videos

See all

Related articles

See all