> Markdown version of [/jobs/ext/2958220-remote](https://www.wearedevelopers.com/jobs/ext/2958220-remote). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Remote - **Company:** MAG 24 LLC - **Location:** New York, NY, United States (Remote available) - **Salary:** $300,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Statistical Hypothesis Testing, Machine Learning, System Testing, Multi-Agent Systems, Information Technology, Deployment Automation, Automation Anywhere, Data Selection - **Published:** September 17, 2026 - **Apply:** https://www.careerjet.com/jobad/us0af198ee9a091b72f27abf5f358f6fa6 ## About the Role * Master's degree in Computer Science, Machine Learning, Artificial Intelligence, or a closely related technical discipline * Strong judgement regarding research-signal quality, data selection, and evaluation design * Experience designing datasets, evaluation frameworks, or QA processes for machine-learning systems * Ability to translate ambiguous operational issues into structured research and evaluation problems * Familiarity with reinforcement-learning environments, agentic systems, or AI-system evaluation * Strong analytical skills and ability to produce concise, actionable technical insights * Proven ability to execute effectively within rapid iteration cycles and high-ambiguity environments * Strong written and verbal communication skills * Collaborative experience across research, product, engineering, and domain teams * Client-facing experience within technical or research-focused environments is advantageous * Experience building internal research or evaluation tooling is beneficial * Contributions to benchmarks, research publications, or open research initiatives are advantageous * Exposure to enterprise AI deployments or forward-deployed research environments is strongly valued ## Description We are sharing a specialised full-time opportunity for experienced technical professionals to operate at the intersection of enterprise AI, applied research, machine-learning evaluation, and real-world AI system performance. Selected professionals will work directly within enterprise AI workflows to identify real-world failure modes, design high-signal datasets and evaluation frameworks, and run rapid experimental cycles that improve system performance. The role combines forward-deployed research, ML-oriented data design, agentic workflow evaluation, technical analysis, and close collaboration across research, product, domain, and enterprise teams., Enterprise AI Research & Failure Analysis * Embed within enterprise AI workflows as a technical research collaborator * Work alongside domain experts and enterprise teams to understand real-world system behaviour * Identify, formalise, and prioritise failure modes emerging from deployed AI systems * Translate operational issues into structured research questions and measurable technical problems * Produce clear analyses of system behaviour, limitations, and opportunities for improvement ML-Oriented Data & Evaluation Design * Design high-signal datasets targeting identified model and system weaknesses * Develop evaluation protocols, quality criteria, and structured assessment frameworks * Apply strong judgement to data selection, evaluation design, and research-signal quality * Identify gaps in existing datasets and evaluation coverage * Structure research workflows to support measurable improvements in model performance Experimentation & Agentic Workflow Evaluation * Run rapid experimental cycles to test hypotheses and quantify system improvements * Develop and benchmark agentic workflows with a focus on robustness, reliability, and scalability * Evaluate AI systems operating across complex enterprise workflows * Analyse experimental results and determine whether observed improvements are meaningful and reproducible * Iterate on datasets, evaluations, and system configurations based on research findings Research Tooling & Cross-Functional Collaboration * Build lightweight tooling to support evaluation, data curation, experimentation, and rapid iteration * Collaborate across research, engineering, product, domain, and enterprise-facing teams * Translate research findings into clear, decision-oriented recommendations * Contribute to research artifacts including reports, benchmarks, evaluation documentation, and technical analyses * Communicate complex findings clearly to both technical and non-technical stakeholders ## Related Videos - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Intelligent Data Selection for Continual Learning of AI Functions](https://www.wearedevelopers.com/videos/367-intelligent-data-selection-for-continual-learning-of-ai-functions) - [How To Test A Ball of Mud](https://www.wearedevelopers.com/videos/173-how-to-test-a-ball-of-mud) - [Mastering AI-Driven Problem Solving in Engineering with Observability](https://www.wearedevelopers.com/videos/994-mastering-ai-driven-problem-solving-in-engineering-with-observability) - [How I Built QA from Scratch in a Scaling Startup - no fluff real life story](https://www.wearedevelopers.com/videos/2041-how-i-built-qa-from-scratch-in-a-scaling-startup-no-fluff-real-life-story) - [Carl Lapierre - Exploring Advanced Patterns in Retrieval-Augmented Generation](https://www.wearedevelopers.com/videos/1235-carl-lapierre-exploring-advanced-patterns-in-retrieval-augmented-generation) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)