> Markdown version of [/jobs/ext/2960128-remote](https://www.wearedevelopers.com/jobs/ext/2960128-remote). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Remote - **Company:** MAG 24 LLC - **Location:** New York, NY, United States (Remote available) - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Data Integrity, Data Systems, Performance Tuning, Software Requirements Analysis - **Published:** September 17, 2026 - **Apply:** https://www.careerjet.com/jobad/usdb1eb9675f6ff928881e8a87b2c4a53b ## About the Role * Strong professional judgement regarding research signal quality and whether findings are ready to support broader conclusions * Experience designing ML-oriented datasets, evaluation frameworks, annotation systems, rubrics, or QA processes * Ability to translate complex and ambiguous real-world system behaviour into structured research and evaluation opportunities * Strong ownership mindset and comfort making decisions in uncertain or rapidly evolving environments * Excellent written and verbal communication skills * Ability to explain technical trade-offs, limitations, evidence quality, and research findings clearly * Proven experience working directly with researchers, technical experts, or domain specialists during project calibration and iteration * Systems-level understanding of model, agent, or AI-system performance * Experience with reinforcement-learning environments, simulators, or feedback-driven training systems is advantageous * Experience improving agentic systems or AI systems operating within real-world workflows is beneficial * Prior work within applied research or production environments with direct impact on deployed systems is advantageous * Experience designing evaluations for complex or real-world tasks is strongly valued * Familiarity with expert incentive design or high-stakes technical research programmes is beneficial ## Description We are sharing a specialised full-time opportunity for experienced technical professionals to operate at the intersection of AI research, machine-learning data systems, evaluation, and real-world model performance. Selected professionals will take ownership of research and evaluation initiatives designed to generate high-quality, defensible research signal and translate that signal into measurable improvements in AI systems. The role combines evaluation design, ML-oriented data development, failure analysis, quality calibration, and close collaboration with researchers, domain experts, and operational teams., Research Evaluation & Signal Quality * Own research and evaluation initiatives from problem framing through data design, quality calibration, and signal validation * Define rigorous approaches for determining whether experimental results provide reliable and defensible research signal * Analyse model and system failures to identify root causes, edge cases, and opportunities for improvement * Evaluate whether datasets, experiments, and conclusions meet appropriate quality thresholds * Act as a quality gate when signal strength, data integrity, or supporting evidence is insufficient ML-Oriented Data & Evaluation Design * Design ML-oriented data systems including task definitions, annotation schemas, rubrics, incentives, and supporting pipelines * Structure data and evaluation workflows around downstream model-performance objectives * Translate ambiguous real-world behaviour into measurable evaluation frameworks and new data categories * Identify gaps in evaluation or dataset coverage and recommend where additional investment or iteration is needed * Develop quality-assurance processes that maintain strong and consistent research standards Failure Analysis & Iterative Model Improvement * Investigate model and system behaviour to identify recurring weaknesses and performance limitations * Iterate rapidly on evaluations, datasets, feedback loops, and quality standards * Use experimental findings to guide improvements in model or agent performance * Determine when research directions should be expanded, revised, paused, or discontinued based on evidence * Maintain a systems-level perspective focused on end-to-end AI performance rather than isolated components Research Collaboration & Technical Communication * Work closely with researchers, domain experts, operators, and cross-functional teams throughout project kickoff, calibration, and iteration * Communicate research findings, trade-offs, limitations, and signal strength clearly to technical and non-technical stakeholders * Translate research progress into credible narratives grounded in evidence * Support alignment between experimental work and real-world system requirements * Contribute strong technical judgement in ambiguous, high-impact research environments ## Related Videos - [Carl Lapierre - Exploring Advanced Patterns in Retrieval-Augmented Generation](https://www.wearedevelopers.com/videos/1235-carl-lapierre-exploring-advanced-patterns-in-retrieval-augmented-generation) - [Practical performance tuning for Serverless Java on AWS](https://www.wearedevelopers.com/videos/2075-practical-performance-tuning-for-serverless-java-on-aws) - [Exploring 5 Key Applications of AI Abundance with Blockchain Assurance](https://www.wearedevelopers.com/videos/971-exploring-5-key-applications-of-ai-abundance-with-blockchain-assurance) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Mastering AI-Driven Problem Solving in Engineering with Observability](https://www.wearedevelopers.com/videos/994-mastering-ai-driven-problem-solving-in-engineering-with-observability) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) ## Related Articles - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)