Research Data Scientist

TALENTHOP LLC
United States
3 days ago
Apply on arc.dev
Prepare application

Role details

Contract type
Permanent contract
Employment type
Part-time (≤ 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Data Analysis Big Data Information Engineering Statistical Hypothesis Testing Python (Programming Language) Machine Learning Natural Language Processing NumPy Tensorflow SQL Databases Feature Engineering
+7 more
Pytorch Large Language Models Model Validation Generative AI Pandas Scikit Learn Information Technology

Job description

The RESEARCH DATA SCIENTIST role in the AI/LLM Delivery Unit centers on leading research-driven AI and machine learning initiatives involving Generative AI, large language models, natural language processing, and model evaluation. This role is critical for designing experiments, developing evaluation methods, analyzing complex datasets, and transforming research into practical AI solutions. It contributes to advancing the organization’s capabilities in AI by collaborating across multiple technical and client-facing teams., * Conduct independent and collaborative research in Generative AI, LLMs, NLP, multimodal AI, and model evaluation.

  • Formulate research questions and develop structured methodologies and experiments.
  • Design and analyze experiments to improve AI/ML models and solutions.
  • Build analytical models, prototypes, and research pipelines using Python and ML frameworks.
  • Develop and implement LLM evaluation frameworks, benchmarks, and datasets.
  • Evaluate models for accuracy, robustness, bias, hallucination, and other performance aspects.
  • Perform model benchmarking, error analysis, and comparative performance evaluation.
  • Work on techniques like RAG, SFT, RLHF/DPO, prompt engineering, fine-tuning, embeddings, and optimization.
  • Identify gaps in models and data, recommending improvements.
  • Collect, clean, analyze, and interpret complex datasets.
  • Perform exploratory data analysis, statistical testing, and error analysis.
  • Develop data-driven insights relevant to AI/ML research.
  • Develop and evaluate datasets, annotation frameworks, and data quality metrics.
  • Collaborate with annotation, data engineering, and AI/ML teams to enhance training and evaluation data.
  • Contribute to research publications, technical reports, patents, benchmarks, and internal documents.
  • Identify and explore new methodologies, models, and evaluation approaches.
  • Engage with stakeholders to present findings and translate technical concepts into actionable recommendations.
  • Participate in client-facing technical discussions as needed.

Requirements

  • Master’s or PhD in Computer Science, AI, Machine Learning, Data Science, Statistics, Mathematics, Computational Science, or related field.
  • Bachelor’s or Master’s degree from premier engineering or research institutions is strongly preferred.
  • 4 to 7 years of hands-on research experience in AI/ML, Data Science, NLP, Generative AI, or related fields.
  • Proven ability to independently formulate research questions, design experiments, analyze results, and communicate findings.
  • Demonstrated research contributions through publications, patents, presentations, or significant projects.
  • Proficiency in Python and SQL.
  • Experience with NumPy, Pandas, Scikit-learn, and preferably PyTorch or TensorFlow.
  • Strong understanding of machine learning algorithms, statistics, experimentation, data analysis, feature engineering, model evaluation, and hypothesis testing.
  • Hands-on experience with LLMs, NLP, Generative AI, and multimodal AI.

Benefits & conditions

Pay Range and Compensation Package:

  • The pay range and compensation package for this role will be determined based on the candidate’s experience, skills, and other relevant factors.

About the company

The organization operates in the data engineering industry with a focus on advancing artificial intelligence responsibly. It addresses the challenge of building trustworthy AI systems by providing essential data, evaluation frameworks, and human expertise. The company offers a range of solutions, platforms, and services that support Generative AI and AI system developers, grounded in over 36 years of delivering high-quality data and measurable outcomes.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on arc.dev
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:01 min

Executing remote data exploration and model training

Mingshen Sun Mingshen Sun · World Congress 2024

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:25 min

Replacing NumPy with cuPy for straightforward GPU acceleration

Paul Graham Paul Graham · World Congress 2025

Videos

See all

Related articles

See all