Remote

MAG 24 LLC
New York, NY, United States
3 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence Data Integrity Data Systems Performance Tuning Software Requirements Analysis

Job description

We are sharing a specialised full-time opportunity for experienced technical professionals to operate at the intersection of AI research, machine-learning data systems, evaluation, and real-world model performance. Selected professionals will take ownership of research and evaluation initiatives designed to generate high-quality, defensible research signal and translate that signal into measurable improvements in AI systems. The role combines evaluation design, ML-oriented data development, failure analysis, quality calibration, and close collaboration with researchers, domain experts, and operational teams., Research Evaluation & Signal Quality

  • Own research and evaluation initiatives from problem framing through data design, quality calibration, and signal validation
  • Define rigorous approaches for determining whether experimental results provide reliable and defensible research signal
  • Analyse model and system failures to identify root causes, edge cases, and opportunities for improvement
  • Evaluate whether datasets, experiments, and conclusions meet appropriate quality thresholds
  • Act as a quality gate when signal strength, data integrity, or supporting evidence is insufficient

ML-Oriented Data & Evaluation Design

  • Design ML-oriented data systems including task definitions, annotation schemas, rubrics, incentives, and supporting pipelines
  • Structure data and evaluation workflows around downstream model-performance objectives
  • Translate ambiguous real-world behaviour into measurable evaluation frameworks and new data categories
  • Identify gaps in evaluation or dataset coverage and recommend where additional investment or iteration is needed
  • Develop quality-assurance processes that maintain strong and consistent research standards

Failure Analysis & Iterative Model Improvement

  • Investigate model and system behaviour to identify recurring weaknesses and performance limitations
  • Iterate rapidly on evaluations, datasets, feedback loops, and quality standards
  • Use experimental findings to guide improvements in model or agent performance
  • Determine when research directions should be expanded, revised, paused, or discontinued based on evidence
  • Maintain a systems-level perspective focused on end-to-end AI performance rather than isolated components

Research Collaboration & Technical Communication

  • Work closely with researchers, domain experts, operators, and cross-functional teams throughout project kickoff, calibration, and iteration
  • Communicate research findings, trade-offs, limitations, and signal strength clearly to technical and non-technical stakeholders
  • Translate research progress into credible narratives grounded in evidence
  • Support alignment between experimental work and real-world system requirements
  • Contribute strong technical judgement in ambiguous, high-impact research environments

Requirements

  • Strong professional judgement regarding research signal quality and whether findings are ready to support broader conclusions
  • Experience designing ML-oriented datasets, evaluation frameworks, annotation systems, rubrics, or QA processes
  • Ability to translate complex and ambiguous real-world system behaviour into structured research and evaluation opportunities
  • Strong ownership mindset and comfort making decisions in uncertain or rapidly evolving environments
  • Excellent written and verbal communication skills
  • Ability to explain technical trade-offs, limitations, evidence quality, and research findings clearly
  • Proven experience working directly with researchers, technical experts, or domain specialists during project calibration and iteration
  • Systems-level understanding of model, agent, or AI-system performance
  • Experience with reinforcement-learning environments, simulators, or feedback-driven training systems is advantageous
  • Experience improving agentic systems or AI systems operating within real-world workflows is beneficial
  • Prior work within applied research or production environments with direct impact on deployed systems is advantageous
  • Experience designing evaluations for complex or real-world tasks is strongly valued
  • Familiarity with expert incentive design or high-stakes technical research programmes is beneficial

Benefits & conditions

Engagement Details

  • Full-time engagement
  • Fully remote
  • Compensation: $600,000-$2,000,000/year
  • Work will span research evaluation, ML-oriented data design, failure analysis, quality calibration, and iterative AI-system improvement
  • Responsibilities may include acting as a quality gate for research claims, datasets, and evaluation results
  • Collaboration will involve researchers, domain experts, operational teams, and other technical stakeholders
  • Project priorities, evaluation frameworks, and research directions may evolve based on experimental findings and system performance
  • Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party

About the Platform This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

2:37 min

Classifying and anonymizing data during system design

Reto Kaeser · LIVE

1:55 min

Maintaining software and data integrity during execution

Christian Wenz Christian Wenz · World Congress 2026 Europe

8:32 min

Benchmarking GitOps engine constraints for extensive multi-cluster environments

Artem Lajko · Europe 2026 Virtual

2:12 min

Navigating technical clarity as a global black belt

Chris Heilmann +2 · LIVE

4:15 min

Bridging operational and analytical systems using formal data contracts

Matthias Niehoff Matthias Niehoff · World Congress 2024

Videos

See all

Related articles

See all