> Markdown version of [/jobs/ext/1916301-data-scientist-for-ai-model-evaluation](https://www.wearedevelopers.com/jobs/ext/1916301-data-scientist-for-ai-model-evaluation). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist for AI Model Evaluation - **Company:** Amazon.com, Inc. - **Location:** Seattle, WA, United States (Remote available) - **Salary:** $124,800.0 - $187,200.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Data Analysis, Data Cleansing, Statistical Hypothesis Testing, Python (Programming Language), NumPy, Jupyter Notebook, Model Validation, Git, Pandas - **Published:** August 4, 2026 - **Apply:** https://www.careerjet.com/job/register/usa271c2acb2aacb16cd5484d532889e56 ## About the Role * MSc or PhD in statistics, data science, or another quantitative STEM field, or equivalent practical experience in a research-heavy analytical domain. * At least 1 year of experience in a research, research-engineering, or heavy data-analysis role. * Deep, hands-on skills in data cleaning, statistical correlation, hypothesis testing, and careful interpretation of results. * Proficiency using Jupyter Notebooks or Google Colab for analysis and reporting. * Working proficiency in Python, including libraries such as pandas and NumPy, and familiarity with Git. * Strong written communication skills, capable of explaining analytical findings to decision-makers. * Experience in AI training, model evaluation, or benchmark/task authoring is preferred. * A perfectionist mindset, high attention to detail, creativity in task design, and the ability to work independently on ambiguous, open-ended problems. * Ability to engage reliably for approximately 35 hours per week. ## Description Help build next-generation agentic evaluation benchmarks for frontier AI models by acting as a ground-truth expert in data science and quantitative analysis. You will design and execute realistic, research-style analysis tasks that test model capabilities, verify statistical claims, and produce clear, reproducible notebooks that drive researcher decisions. Tasks typically require one to two days of continuous, focused effort and span data cleaning, statistical analysis, interpretation, and written reporting. You will operate in a close feedback loop with researchers to pinpoint where advanced models fall short on rigorous analytical work. Key Responsibilities * Design complex, realistic data-analysis tasks that simulate actual research work, including cleaning messy datasets, defining fair comparisons between methods, and specifying success criteria. * Author reproducible analyses and reports in Jupyter Notebooks or Google Colab that a researcher can follow and act on. * Construct comparisons between analytical approaches, for example comparing two anomaly-detection algorithms on a dataset, calculating correlations, and recommending a preferred method backed by spot checks. * Perform manual spot checks and statistical validation, interpret results carefully, and summarize findings clearly enough to inform research decisions. * Evaluate how models perform on your tasks, confirming whether reported statistics and conclusions hold up under scrutiny. * Coordinate with researchers and other subject-matter experts to align evaluation standards and maintain consistency across tasks. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Getting to Know Your Legacy (System) with AI-Driven Software Archeology](https://www.wearedevelopers.com/videos/1437-getting-to-know-your-legacy-system-with-ai-driven-software-archeology) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [How to start an AI project for a good cause and boost your career](https://www.wearedevelopers.com/magazine/15-how-to-start-an-ai-project-for-a-good-cause-and-boost-your-career) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)