> Markdown version of [/jobs/ext/2640257-data-scientist](https://www.wearedevelopers.com/jobs/ext/2640257-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist - **Company:** SumasEdge Corporation - **Location:** Louisville, KY, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Data Analysis, Python (Programming Language), Machine Learning, Software Safety, Large Language Models, Model Validation, Generative AI, Information Technology - **Published:** August 28, 2026 - **Apply:** https://www.dice.com/job-detail/f21ad5fa-e956-4e6b-b5a1-97e116efeb09 ## About the Role * Master's degree or PhD in Data Science, Statistics, Computer Science, Machine Learning, or related field, or equivalent experience. * 3+ years of experience in data science, machine learning, analytics, or AI model evaluation. * Experience analyzing voice-based language data. * Strong experience with Python and data science tooling. * Experience designing experiments and evaluating machine learning models. * Knowledge of statistical analysis, sampling methodologies, and model validation techniques. * Experience with Generative AI, LLMs, or conversational AI systems. * Familiarity with AI safety, Responsible AI, fairness, bias, and governance frameworks. * Experience developing automated evaluation systems. * ## Description We are seeking a Data Scientist to help build and maintain a robust evaluation framework for conversational AI systems. This role will focus on developing automated AI quality measurement solutions, including LLM-as-a-Judge systems, to assess hallucination rates, intent accuracy, transcript quality, and responsible AI metrics at scale. As a key member of our AI Quality team, you will work closely with annotation specialists, product teams, and governance stakeholders to ensure our AI experiences are accurate, safe, compliant, and customer-centric. * Design and develop LLM-as-a-Judge evaluation frameworks for conversational AI systems. * Build automated monitoring solutions to evaluate millions of voice and chat-based customer interactions. * Collaborate with user experience teams to identify core metrics and evaluation criteria. * Validate model performance against human-labeled datasets using statistical best practices. * Measure and report precision, recall, false-positive rates, and false-negative rates. * Develop statistical sampling methodologies and quality measurement frameworks. * Continuously calibrate and improve evaluation models as production AI systems evolve. * Analyze hallucination rates, intent classification accuracy, transcript quality (WER), guardrail effectiveness, vulnerability to jailbreak attempts, and fairness metrics. * Create dashboards and reporting to support governance, legal, compliance, and Responsible AI reviews. * Collaborate with annotation teams to improve label quality and gold-standard datasets. * Support future multilingual AI evaluation initiatives. ## Related Videos - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Your imaginations is (no longer) the limit: how Generative AI empowers people to be creative](https://www.wearedevelopers.com/videos/741-your-imaginations-is-no-longer-the-limit-how-generative-ai-empowers-people-to-be-creative) - [Psychological Safety in Software Engineering - Jenny-Margrethe Vej & Alexandra Hou Aldershaab](https://www.wearedevelopers.com/videos/2142-psychological-safety-in-software-engineering-jenny-margrethe-vej-alexandra-hou-aldershaab) - [Beyond Dashboards: Fixing Text-to-SQL with Semantic RAG](https://www.wearedevelopers.com/videos/2036-beyond-dashboards-fixing-text-to-sql-with-semantic-rag) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) - [The shadows that follow the AI generative models](https://www.wearedevelopers.com/videos/624-the-shadows-that-follow-the-ai-generative-models) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)