> Markdown version of [/jobs/ext/1469425-lead-data-scientist](https://www.wearedevelopers.com/jobs/ext/1469425-lead-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Data Scientist - **Company:** octaves llc - **Location:** Roanoke, VA, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** A/B Testing, Artificial Intelligence, Airflow, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, JIRA, Big Data, Continuous Integration, Extract Transform Load (ETL), Database Queries, Distributed Computing Environment, Monitoring of Systems, Python (Programming Language), Machine Learning, NumPy, Performance Tuning, Power BI, Tensorflow, Prometheus, SQL Databases, Tableau (Software), Management of Software Versions, Datadog, Data Processing, Feature Engineering, Pytorch, Flask (Web Framework), Delivery Pipeline, Large Language Models, Snowflake, Prompt Engineering, Apache Spark, Deep Learning, Generative AI, Fastapi, Pandas, Build Management, Pytest, Scikit Learn, Kubernetes, Information Technology, HuggingFace, Data Analytics, Xgboost, Data Management, Machine Learning Operations, Virtual Agents, Functional Programming, Api Design, Cloudwatch, Document Classification, GPT, Software Version Control, GXP, Docker, Unsupervised Learning, Databricks - **Published:** July 28, 2026 - **Apply:** https://www.dice.com/job-detail/dd5badbd-cdca-418a-baa0-d4a819e2f770 ## About the Role Python (5+ years): Production-level experience with Pandas, NumPy, scikit-learn, XGBoost, TensorFlow/PyTorch, Hugging Face Transformers, FastAPI/Flask, MLflow, and pytest SQL: Advanced proficiency with complex queries, window functions, and optimization Machine Learning & NLP: Strong foundation in supervised/unsupervised learning, deep learning, document understanding, text classification, and semantic analysis Generative AI & LLMs: Hands-on experience with foundation models (GPT, Claude, Llama), prompt engineering, RAG architectures, and vector databases (Pinecone, Weaviate, Chroma) MLOps & ModelOps: End-to-end experience with ML pipelines, experiment tracking (MLflow, W&B), model versioning, feature stores, drift detection, CI/CD for ML, and Docker containerization LLM Evaluation: Experience with evaluation frameworks (RAGAS, DeepEval), custom metrics, benchmark datasets, and human-in-the-loop validation Cloud & AWS: Experience with AWS services including SageMaker, Bedrock, S3, Lambda, EC2, and CloudWatch Statistics & Experimentation: Strong foundation in statistics, A/B testing, causal inference, and experimental design Visualization: Proficiency with Tableau, Power BI, or Python visualization libraries, * 7+ years in data science, ML engineering, or related roles * 3+ years building NLP/generative AI applications and implementing MLOps in production * Bachelor's or Master's degree in Data Science, Computer Science, Statistics, or related field (PhD preferred) * Track record of deploying ML systems processing large-scale datasets with proper monitoring and governance Preferred Qualifications * Experience with agentic AI frameworks (LangGraph, LangChain, AutoGen, CrewAI) * Knowledge of Life Sciences/regulated industries (FDA, EMA, ISO, GxP) and compliance management systems * Familiarity with big data tools (Spark, Databricks, Snowflake), orchestration (Airflow, Kubeflow), and monitoring tools (Datadog, Prometheus) * Experience with LLM fine-tuning, document processing libraries, multi-modal AI, or distributed training * Understanding of ML governance, bias detection, model risk management, and data privacy regulations (GDPR, CCPA, HIPAA) * Experience working in agile environments with Jira * AWS ML certifications or similar credentials Key Competencies * Strong communication skills explaining complex models to technical and non-technical audiences * Ability to work independently and collaboratively in fast-paced environments * Proven ability to convert POCs into production-grade solutions * Understanding of ethical AI and building trustworthy, explainable systems for regulated environments ## Description Octave's ETQ division is seeking a hands-on Data Scientist to build predictive models, implement Generative AI and Agentic AI features, and architect data-driven solutions for our document-based compliance management platform. This role requires a technical expert who can develop, deploy, and maintain ML systems in production environments. * Build and deploy Generative AI features using foundation models (AWS Bedrock, OpenAI, Anthropic Claude) and RAG architectures with vector databases for compliance document understanding * Design agentic AI systems that autonomously handle compliance workflows, document review, regulatory mapping, and multi-step reasoning tasks * Implement comprehensive LLM evaluation frameworks with automated pipelines, custom metrics, benchmark datasets, and safety guardrails ensuring regulatory compliance * Build end-to-end MLOps pipelines for model training, deployment, monitoring, versioning, and automated retraining with drift detection * Develop predictive models for compliance risk scoring, regulatory change impact, anomaly detection, and time-series forecasting * Write production-quality Python code for data processing, feature engineering, API development (FastAPI/Flask), and ETL/ELT workflows * Lead A/B experiments and product analytics to measure AI feature impact and drive data-driven decision-making * Create explainability frameworks (SHAP/LIME) and monitoring dashboards ensuring transparency and regulatory adherence * Collaborate with cross-functional teams to translate business needs into ML solutions and communicate insights to stakeholders ## Related Videos - [Blueprints for Success: Steering a Global Data & AI Architecture](https://www.wearedevelopers.com/videos/1577-blueprints-for-success-steering-a-global-data-ai-architecture) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Collaboration Quantified: Lessons from Open Source Developer Networks](https://www.wearedevelopers.com/videos/1422-collaboration-quantified-lessons-from-open-source-developer-networks) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [Integrate your Cognitive Assistant with 3rd-party DBs and software](https://www.wearedevelopers.com/videos/249-integrate-your-cognitive-assistant-with-3rd-party-dbs-and-software) - [Streaming AI Responses in Real-Time with SSE in Next.js & NestJS](https://www.wearedevelopers.com/videos/1630-streaming-ai-responses-in-real-time-with-sse-in-next-js-nestjs) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)