Deep Research Task Evaluator

Human Union Data, Inc.
United States
13 days ago

Role details

Contract type
Permanent contract
Employment type
Part-time (≤ 32 hours)
Working hours
Regular working hours
Languages
English
Job source

Tech stack

Artificial Intelligence Model Validation

Job description

Deep Research Task Evaluator is a remote review track for evaluating AI outputs across deep research task research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct method so the modeling team can train on it.

Why this role matters

Deep Research Task research review models live or die on whether their derivations actually hold up under scrutiny. AuraOne uses scientific specialists to grade outputs the way a peer reviewer would - checking assumptions, reproducing key steps, and capturing the right method alongside the wrong one., * Review AI outputs against current deep research task research review methods, conventions, and prior work for Deep Research Task Evaluator assignments.

  • Reproduce or sanity-check key derivations, calculations, or experimental claims.
  • Flag dimensional, methodological, and citation errors with structured severity tags.
  • Capture the corrected reasoning or worked example so the modeling team can train on it.
  • Adjudicate disputed answers against textbooks, papers, or community standards.
  • Maintain reviewer-quality scores in inter-rater calibration cycles., * Reproduce a deep research task research review derivation from a model output and flag any algebraic or dimensional errors.
  • Grade a model’s literature summary against the cited papers and rate the citation quality.
  • Adjudicate a disputed answer between two reviewers using textbook methods.
  • Audit a 25-row batch for rubric consistency and report drift to the program lead.

Requirements

  • Graduate-level training or equivalent applied experience in deep research task research review or a closely related field for Deep Research Task Evaluator work.
  • Hands-on experience publishing, teaching, or advising on the topic at a professional level.
  • Comfort applying multi-page rubrics consistently across long batches.
  • Clear written reasoning that cites methods, papers, or worked examples.
  • Reliable async availability for at least 10 hours per week., * PhD, postdoc, or industry research experience in the topic area.
  • Prior work reviewing AI-assisted research tooling and its failure modes.
  • Multilingual fluency for non-English papers and corpora., * Scientific reasoning
  • Method validation
  • Citation review
  • Quantitative analysis
  • Deep Research Task research review
  • Web research
  • Source grounding
  • Browser automation
  • Deep
  • Research

Work model

Remote - US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.

Benefits & conditions

Hourly rate confirmed after the interview process.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Assisting manual license research with artificial intelligence tools

Uwe Korn Uwe Korn · WWC Europe 2026

3:46 min

Core terminology and audiences for interpretable artificial intelligence

Karol Przystalski · LIVE

5:08 min

Validating requests and responses using data transfer objects

Roman Alexis Anastasini · WWC 2021

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

3:21 min

Automating complete quality assurance pipelines with artificial intelligence

Evelyn Haslinger · LIVE

56 sec

The negligible impact of AI model size on security

Julian Totzek-Hallhuber Julian Totzek-Hallhuber · WWC Europe 2026

Videos

See all

Related articles

See all