Data Scientist

Datavault AI Inc.
Atlanta, GA, United States
about 2 months ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Databases Customer Data Management Database Queries Statistical Hypothesis Testing Python (Programming Language) PostgreSQL Machine Learning NumPy Tensorflow Pytorch Large Language Models
+6 more
Pandas Scikit Learn Information Technology Production Code Power Analysis (Cryptography) Machine Learning Operations

Job description

We’re looking for a Data Scientist to drive measurable improvement of our AI systems - including a multi-agent LLM pipeline that profiles, classifies, and values customer data assets, and a classification service that builds our reference dataset from public sources. You’ll own the evaluation strategy, ground-truth corpus design, and statistical rigor that turns “the agent feels better” into “the agent is measurably 18% more accurate at industry classification on our latest corpus.” This is a hands-on, high-ownership role where you’ll be the technical authority on what “good” looks like for our AI outputs., * Design, build, and maintain evaluation frameworks for our multi-agent LLM pipelines covering classification, PII detection, valuation, retrieval, segmentation, and synthesis - with regression-detection rigor.

  • Curate, expand, and version synthetic and real-world test corpora that exercise our AI pipelines end-to-end across 20+ industry verticals.
  • Quantify model performance: precision, recall, calibration, inter-rater agreement against human-verified ground truth, drift detection across releases.
  • Partner with engineering to design prompt experiments, agent variants, and structured-output schema iterations; report results with statistical confidence intervals - not anecdotes.

  • Improve vector-search comparable retrieval: embedding model selection, retrieval evaluation (recall@k, MRR), taxonomy refinement, classification accuracy uplift.
  • Evaluate prompt strategies, tool-use patterns, and routing logic; recommend model-tier choices backed by cost/accuracy data.
  • Profile production traces to identify failure modes (hallucinated outputs, mis-classifications, missed PII), then design experiments to fix them.
  • Work cross-functionally with engineering, product, and domain experts to translate fuzzy product goals (“the analysis should feel insightful”) into quantitative success metrics.

  • Communicate findings through written reports, dashboards, and decision memos that executive leadership can act on.

Requirements

  • Bachelor’s degree in Computer Science, Statistics, Data Science, Machine Learning, or related quantitative discipline, or equivalent professional experience. Master’s or PhD preferred.
  • 3+ years of professional data science, ML engineering, or AI evaluation experience shipping models or AI systems to production.
  • Strong Python skills (Pandas, NumPy, scikit-learn, PyTorch or TensorFlow), with comfort writing production-quality code that engineers will run in CI.
  • Proven experience evaluating LLM-based systems: prompt experimentation, structured-output validation, hallucination detection, retrieval evaluation, judge-LLM patterns.
  • Solid grasp of classical statistics: hypothesis testing, confidence intervals, sample-size calculation, power analysis, calibration.
  • SQL proficiency for ad-hoc analysis on PostgreSQL; comfortable with embedded analytical databases for offline evaluation.
  • Experience with vector databases and embedding models.

  • Ability to translate business goals into measurable evaluation criteria, and willingness to push back when a “metric” doesn’t measure what stakeholders thin

Benefits & conditions

  • Competitive salary and benefits package.
  • A fast-paced, high-impact work environment.
  • Opportunity to work closely with executive leadership.

  • The chance to work with cutting-edge technologies and make a significant impact.
  • A culture of innovation, ownership, and growth.

About the company

About Us: Datavault AI, along with its event-technology subsidiary Event Citadel (formerly CompuSystems), operates across a diverse portfolio of technology and service divisions.

Datavault AI Inc. delivers high-performance computing software, Web 3.0 data-management solutions, and advanced audio technologies to a broad range of industries.

Event Citadel (formerly CompuSystems), founded in 1976, is a trusted provider of end-to-end event technology solutions, offering registration, ticketing, lead retrieval, and attendee-engagement services for events of all sizes across trade, association, corporate, and government markets.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou ¡ Coffee With Developers

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell ¡ LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel ¡ World Congress 2024

2:35 min

Preventing remote code execution in PyTorch models

Balåzs Kiss ¡ World Congress 2023

1:50 min

Introduction to the speaker and data science background

Bas Geerdink ¡ LIVE

1:25 min

Replacing NumPy with cuPy for straightforward GPU acceleration

Paul Graham Paul Graham ¡ World Congress 2025

Videos

See all

Related articles

See all