> Markdown version of [/jobs/ext/131516-senior-data-scientist](https://www.wearedevelopers.com/jobs/ext/131516-senior-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Scientist - **Company:** Zywave, Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** A/B Testing, Artificial Intelligence, Data Analysis, Automation of Tests, Data Infrastructure, Extract Transform Load (ETL), Data Warehousing, Monitoring of Systems, Statistical Hypothesis Testing, Python (Programming Language), Machine Learning, SQL Databases, Workflow Management Systems, Feature Engineering, Large Language Models, Grafana, Prompt Engineering, Model Validation, Generative AI, Pandas, Build Management, Scikit Learn, Low Latency, Power Analysis (Cryptography), Machine Learning Operations, Tools for Reporting, Software Version Control, Data Pipelines, Unsupervised Learning - **Published:** May 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=b2effcd23309d152 ## About the Role * Strong foundation in statistical methods, experimental design, and causal inference-you understand the math behind A/B testing and can design experiments that answer complex questions. * Proven experience building evaluation frameworks or quality measurement systems for ML/AI products in production environments. * Deep expertise in data workflows, including ETL/ELT, data modeling, and pipeline orchestration tools. * Solid machine learning fundamentals across supervised/unsupervised learning, model evaluation, and feature engineering. * Proficiency in Python and the data science stack (pandas, scikit-learn, SQL, visualization libraries). * Experience with modern data platforms and tools (data warehouses, workflow orchestration, version control). * Strong communication skills-you can translate complex statistical concepts for diverse audiences and drive alignment on technical decisions. * Passion for quality and a methodical approach to measuring what matters. * (Preferred) Experience with generative AI evaluation, LLM observability tools, or prompt engineering. ## Description At Zywave, we believe in building AI systems that are reliable, measurable, and continuously improving. The Senior Data Scientist will partner directly with our VP of AI Engineering and Data Science to implement our evaluation framework for agentic AI models. This role is deeply technical, requiring expertise in statistical rigor, and production ML workflows. You'll establish the methodologies and infrastructure that ensure our AI systems meet quality standards and drive measurable business impact. What you will do: Evaluation Framework Design & Implementation * Design and build comprehensive evaluation frameworks for agentic AI models, including benchmark creation, metric definition, and success criteria. * Establish evaluation pipelines that integrate seamlessly into our ML development lifecycle. * Create automated testing suites that assess model performance across multiple dimensions (accuracy, latency, cost, safety, user experience). * Develop methodologies for evaluating complex agent behaviors, multi-turn interactions, and reasoning capabilities. Experimentation & A/B Testing * Design and execute rigorous A/B tests and multivariate experiments to measure model performance in production. * Build statistical frameworks for experiment analysis, including power analysis, significance testing, and causal inference. * Partner with engineering teams to implement experimentation infrastructure and ensure proper randomization and isolation. * Communicate experiment results and recommendations to technical and non-technical stakeholders. Data Science & ML Operations * Apply classical machine learning techniques to support model evaluation, feature engineering, and performance prediction. * Build data pipelines that support evaluation workflows, from data collection through metric computation. * Develop monitoring systems to detect model degradation, drift, and anomalies in production. * Create dashboards and reporting tools that provide visibility into model performance and experimentation results. Strategic Partnership & Leadership * Collaborate closely with the VP of AI Engineering and Data Science to shape our AI quality strategy. * Influence technical decisions through data-driven insights and rigorous analysis. * Establish best practices and standards for evaluation across the organization. * Mentor other data scientists and engineers on experimentation methodologies and statistical principles. Generative AI & LLM Evaluation (Bonus) * Design evaluation approaches specific to generative AI systems (prompt quality, hallucination detection, output consistency). * Develop human-in-the-loop evaluation workflows and LLM-as-judge frameworks. * Stay current with emerging evaluation methodologies in the rapidly evolving GenAI landscape. ## Related Videos - [Bringing the power of AI to your application.](https://www.wearedevelopers.com/videos/1010-bringing-the-power-of-ai-to-your-application) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) - [Beyond Autocomplete: Local AI Code Completion Demystified](https://www.wearedevelopers.com/videos/961-beyond-autocomplete-local-ai-code-completion-demystified) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)