Sr Data Scientist GenAI

Select Minds LLC
Dallas, TX, United States
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Shift work
Job source

Tech stack

Encodings Computational Linguistics Information Retrieval Python (Programming Language) Machine Learning Open Source Technology Tensorflow Search Technologies Software Engineering Data Streaming Data Ingestion Pytorch
+8 more
Retrieval-Augmented Generation Large Language Models Prompt Engineering Generative AI Information Technology HuggingFace Machine Learning Operations Software Version Control

Job description

Sr Data Scientist (NLP / LLM / Generative AI) Location: Dallas, TX Roles & Responsibilities :

  • Design, build, fine-tune, and deploy LLMs, transformer-based NLP models, and GenAI solutions for both batch and real-time/streaming contexts.
  • Own all major components of ML pipelines: data ingestion, cleaning, pre-processing (structured & unstructured), embedding, search & retrieval, prompt engineering, RAG (Retrieval-Augmented Generation).
  • Collaborate closely with ML Engineers, MLOps, software engineering, product, compliance, legal etc., to move models from prototype to production-ensuring reliability, scalability, monitoring, and maintainability.
  • Define and implement evaluation frameworks: accuracy, bias, fairness, hallucination, consistency, latency; run UAT, stress-tests, drift detection.
  • Optimize models and pipelines for performance, cost, and efficiency.
  • Ensure best practices in model development: version control, repeatability, documentation, governance, and ethical AI use.
  • Mentor more junior data scientists; help build team skills in NLP, GenAI practices, prompt engineering, fine-tuning.
  • Identify new use cases; prototype innovations in GenAI/NLP; keep up with latest research and open source developments, decide what to adopt.

Requirements

  • 10+ years of experience in data science / ML, with substantial work in NLP, LLMs, or Generative AI.
  • Deep hands-on experience in Python, using frameworks like PyTorch, TensorFlow, HuggingFace etc.
  • Proven track record building transformer/NLP / LLM models; experience with fine-tuning, prompt engineering.
  • Solid experience with information retrieval / search: keyword + semantic search, embeddings, vector databases.
  • Experience working in production / deploying models (batch and streaming), working with MLOps practices.
  • Strong algorithmic / statistical / mathematical fundamentals. Ability to reason about model behaviour, bias, uncertainty.
  • Good communicator: able to translate complex technical detail to business / non-technical stakeholders. Nice to Have:

  • Master’s in Computer Science, Computational Linguistics, Statistics, Machine Learning or related field.
  • Experience with multimodal models (vision + text) or emerging LLMs and agent-based systems.
  • Experience with open source LLMs & toolkits; familiarity with LangChain or similar frameworks.
  • Prior experience in regulated environments (finance, risk, legal, compliance) with strong governance, privacy requirements.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.wayup.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

3:17 min

Optimizing character encoding with Kim variable byte encoding

Douglas Crockford Douglas Crockford · WWC 2024

1:39 min

Fundamentals of tensors and the TensorFlow library

Håkan Silfvernagel · LIVE

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

4:41 min

Replacing PyTorch with ONNX runtime for AWS Lambda deployments

Marek Suppa · LIVE

2:52 min

Scaling generative AI use cases across large enterprises

Mike Butcher Mike Butcher +3 · WWC 2024

Videos

See all

Related articles

See all