> Markdown version of [/jobs/ext/1282578-data-scientist](https://www.wearedevelopers.com/jobs/ext/1282578-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist - **Company:** Datavault AI Inc. - **Location:** Atlanta, GA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Databases, Customer Data Management, Database Queries, Statistical Hypothesis Testing, Python (Programming Language), PostgreSQL, Machine Learning, NumPy, Tensorflow, Pytorch, Large Language Models, Pandas, Scikit Learn, Information Technology, Production Code, Power Analysis (Cryptography), Machine Learning Operations - **Published:** July 15, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=d3c1c32cf3396037 ## About the Role * Bachelor's degree in Computer Science, Statistics, Data Science, Machine Learning, or related quantitative discipline, or equivalent professional experience. Master's or PhD preferred. * 3+ years of professional data science, ML engineering, or AI evaluation experience shipping models or AI systems to production. * Strong Python skills (Pandas, NumPy, scikit-learn, PyTorch or TensorFlow), with comfort writing production-quality code that engineers will run in CI. * Proven experience evaluating LLM-based systems: prompt experimentation, structured-output validation, hallucination detection, retrieval evaluation, judge-LLM patterns. * Solid grasp of classical statistics: hypothesis testing, confidence intervals, sample-size calculation, power analysis, calibration. * SQL proficiency for ad-hoc analysis on PostgreSQL; comfortable with embedded analytical databases for offline evaluation. * Experience with vector databases and embedding models. * Ability to translate business goals into measurable evaluation criteria, and willingness to push back when a "metric" doesn't measure what stakeholders thin ## Description We're looking for a Data Scientist to drive measurable improvement of our AI systems - including a multi-agent LLM pipeline that profiles, classifies, and values customer data assets, and a classification service that builds our reference dataset from public sources. You'll own the evaluation strategy, ground-truth corpus design, and statistical rigor that turns "the agent feels better" into "the agent is measurably 18% more accurate at industry classification on our latest corpus." This is a hands-on, high-ownership role where you'll be the technical authority on what "good" looks like for our AI outputs., * Design, build, and maintain evaluation frameworks for our multi-agent LLM pipelines covering classification, PII detection, valuation, retrieval, segmentation, and synthesis - with regression-detection rigor. * Curate, expand, and version synthetic and real-world test corpora that exercise our AI pipelines end-to-end across 20+ industry verticals. * Quantify model performance: precision, recall, calibration, inter-rater agreement against human-verified ground truth, drift detection across releases. * Partner with engineering to design prompt experiments, agent variants, and structured-output schema iterations; report results with statistical confidence intervals - not anecdotes. * Improve vector-search comparable retrieval: embedding model selection, retrieval evaluation (recall@k, MRR), taxonomy refinement, classification accuracy uplift. * Evaluate prompt strategies, tool-use patterns, and routing logic; recommend model-tier choices backed by cost/accuracy data. * Profile production traces to identify failure modes (hallucinated outputs, mis-classifications, missed PII), then design experiments to fix them. * Work cross-functionally with engineering, product, and domain experts to translate fuzzy product goals ("the analysis should feel insightful") into quantitative success metrics. * Communicate findings through written reports, dashboards, and decision memos that executive leadership can act on. ## Related Videos - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) - [How to implement convenient Python bindings to C++](https://www.wearedevelopers.com/videos/618-how-to-implement-convenient-python-bindings-to-c) ## Related Articles - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Dev Digest 233: Strip AI Watermarks, Hacked Homelabs & Designing with Code](https://www.wearedevelopers.com/magazine/754-dev-digest-233-strip-ai-watermarks-hacked-homelabs-designing-with-code) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)