Data Science Expert - AI Evaluation

Mercor, Inc.
New York, NY, United States
5 days ago
Apply on www.juju.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$249,600.0
Working hours
Regular working hours
Job source

Tech stack

AI Evaluation A/B Testing Artificial Intelligence Python (Programming Language) SQL Databases

Job description

  • Design precise, task-specific grading criteria for real-world data science deliverables such as analyses, models, dashboards, and experiment readouts.
  • Score AI-generated and human work samples against criteria with detailed written justifications for every score.
  • Apply consistent, evidence-based judgment to ensure scores are reproducible and defensible.
  • Incorporate structured feedback from senior reviewers and iterate quickly on your work.
  • Work independently and asynchronously to meet deadlines while improving AI model performance.

Requirements

Must-Have

  • 5+ years of professional data science experience in industry.
  • Background in business operations, product, or growth data science at top-tier technology companies.
  • Deep fluency in experiment design and A/B testing, metric definition, SQL/Python analysis, and communicating findings to executive stakeholders.
  • Exceptionally strong written communication.
  • Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers.

Preferred

  • Prior experience with AI training, evaluation, or human-data projects.

About the company

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D’Angelo, Larry Summers, and Jack Dorsey.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

54 sec

Generating multiple hook options for outreach A/B testing

Leandro Gomes da Silva Leandro Gomes da Silva · World Congress 2025

6:06 min

Evaluating AI outputs practically using the DeepEval framework

Sebastian Messingfeld Sebastian Messingfeld · World Congress 2026 Europe

1:14 min

Evolution of distributed SQL database architectures

Wei Hu Wei Hu · World Congress 2024

6:49 min

Developing specialized and personalized evaluation benchmarks for AI systems

Andreas Blattmann Andreas Blattmann +4 · World Congress 2025

56 sec

Performing local A/B testing across multiple AI agents

Julia Kasper · Coffee With Developers

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all