Remote
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
We are sharing a specialised full-time opportunity for experienced technical professionals to operate at the intersection of AI research, machine-learning data systems, evaluation, and real-world model performance. Selected professionals will take ownership of research and evaluation initiatives designed to generate high-quality, defensible research signal and translate that signal into measurable improvements in AI systems. The role combines evaluation design, ML-oriented data development, failure analysis, quality calibration, and close collaboration with researchers, domain experts, and operational teams., Research Evaluation & Signal Quality
- Own research and evaluation initiatives from problem framing through data design, quality calibration, and signal validation
- Define rigorous approaches for determining whether experimental results provide reliable and defensible research signal
- Analyse model and system failures to identify root causes, edge cases, and opportunities for improvement
- Evaluate whether datasets, experiments, and conclusions meet appropriate quality thresholds
- Act as a quality gate when signal strength, data integrity, or supporting evidence is insufficient
ML-Oriented Data & Evaluation Design
- Design ML-oriented data systems including task definitions, annotation schemas, rubrics, incentives, and supporting pipelines
- Structure data and evaluation workflows around downstream model-performance objectives
- Translate ambiguous real-world behaviour into measurable evaluation frameworks and new data categories
- Identify gaps in evaluation or dataset coverage and recommend where additional investment or iteration is needed
- Develop quality-assurance processes that maintain strong and consistent research standards
Failure Analysis & Iterative Model Improvement
- Investigate model and system behaviour to identify recurring weaknesses and performance limitations
- Iterate rapidly on evaluations, datasets, feedback loops, and quality standards
- Use experimental findings to guide improvements in model or agent performance
- Determine when research directions should be expanded, revised, paused, or discontinued based on evidence
- Maintain a systems-level perspective focused on end-to-end AI performance rather than isolated components
Research Collaboration & Technical Communication
- Work closely with researchers, domain experts, operators, and cross-functional teams throughout project kickoff, calibration, and iteration
- Communicate research findings, trade-offs, limitations, and signal strength clearly to technical and non-technical stakeholders
- Translate research progress into credible narratives grounded in evidence
- Support alignment between experimental work and real-world system requirements
- Contribute strong technical judgement in ambiguous, high-impact research environments
Requirements
- Strong professional judgement regarding research signal quality and whether findings are ready to support broader conclusions
- Experience designing ML-oriented datasets, evaluation frameworks, annotation systems, rubrics, or QA processes
- Ability to translate complex and ambiguous real-world system behaviour into structured research and evaluation opportunities
- Strong ownership mindset and comfort making decisions in uncertain or rapidly evolving environments
- Excellent written and verbal communication skills
- Ability to explain technical trade-offs, limitations, evidence quality, and research findings clearly
- Proven experience working directly with researchers, technical experts, or domain specialists during project calibration and iteration
- Systems-level understanding of model, agent, or AI-system performance
- Experience with reinforcement-learning environments, simulators, or feedback-driven training systems is advantageous
- Experience improving agentic systems or AI systems operating within real-world workflows is beneficial
- Prior work within applied research or production environments with direct impact on deployed systems is advantageous
- Experience designing evaluations for complex or real-world tasks is strongly valued
- Familiarity with expert incentive design or high-stakes technical research programmes is beneficial
Benefits & conditions
Engagement Details
- Full-time engagement
- Fully remote
- Compensation: $600,000-$2,000,000/year
- Work will span research evaluation, ML-oriented data design, failure analysis, quality calibration, and iterative AI-system improvement
- Responsibilities may include acting as a quality gate for research claims, datasets, and evaluation results
- Collaboration will involve researchers, domain experts, operational teams, and other technical stakeholders
- Project priorities, evaluation frameworks, and research directions may evolve based on experimental findings and system performance
- Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
About the Platform This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Navigating the AI Shift
How to Become an AI Engineer
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
MLOps – What’s the deal behind it?