> Markdown version of [/jobs/ext/3548826-data-science-expert-ai-evaluation](https://www.wearedevelopers.com/jobs/ext/3548826-data-science-expert-ai-evaluation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Science Expert - AI Evaluation - **Company:** Mercor, Inc. - **Location:** New York, NY, United States (Remote available) - **Experience:** Expert - **Salary:** $249,600.0 - **Contract:** Permanent contract - **Skills:** AI Evaluation, A/B Testing, Artificial Intelligence, Python (Programming Language), SQL Databases - **Published:** October 1, 2026 - **Apply:** https://www.juju.com/job/16_2495a8be ## About the Role Must-Have * 5+ years of professional data science experience in industry. * Background in business operations, product, or growth data science at top-tier technology companies. * Deep fluency in experiment design and A/B testing, metric definition, SQL/Python analysis, and communicating findings to executive stakeholders. * Exceptionally strong written communication. * Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers. Preferred * Prior experience with AI training, evaluation, or human-data projects. ## Description * Design precise, task-specific grading criteria for real-world data science deliverables such as analyses, models, dashboards, and experiment readouts. * Score AI-generated and human work samples against criteria with detailed written justifications for every score. * Apply consistent, evidence-based judgment to ensure scores are reproducible and defensible. * Incorporate structured feedback from senior reviewers and iterate quickly on your work. * Work independently and asynchronously to meet deadlines while improving AI model performance. ## Related Videos - [Beyond the Benchmark: How to Evaluate AI Agents in the Real World](https://www.wearedevelopers.com/videos/100269-beyond-the-benchmark-how-to-evaluate-ai-agents-in-the-real-world) - [AI for decision-making in Tech Recruiting](https://www.wearedevelopers.com/videos/1074-ai-for-decision-making-in-tech-recruiting) - [Bringing the power of AI to your application.](https://www.wearedevelopers.com/videos/1010-bringing-the-power-of-ai-to-your-application) - [Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases](https://www.wearedevelopers.com/videos/1146-fault-tolerance-and-consistency-at-scale-harnessing-the-power-of-distributed-sql-databases) - [Testing AI Agents: Automated Evaluation for Chatbots & RAG Systems](https://www.wearedevelopers.com/videos/100300-testing-ai-agents-automated-evaluation-for-chatbots-rag-systems) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) ## Related Articles - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [How to start an AI project for a good cause and boost your career](https://www.wearedevelopers.com/magazine/15-how-to-start-an-ai-project-for-a-good-cause-and-boost-your-career) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it)