Data Scientist

Rite Pros
Portland, ME, United States
8 days ago
Apply on ritepros.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

A/B Testing Adaptable Database Systems Application Programming Interfaces (APIs) Artificial Intelligence Airflow Amazon Web Services Business Analytics Applications Data Analysis Application Performance Management Automation of Tests Microsoft Azure BigQuery
+70 more
Business Software Cloud Computing Cloud Database Code Review Computer Programming Databases Continuous Integration Extract Transform Load (ETL) Data Structures Data Visualization Relational Databases Software Debugging Amazon DynamoDB Statistical Hypothesis Testing Python (Programming Language) PostgreSQL Machine Learning Enterprise Messaging Systems Meta-Data Management Microsoft Message Queuing Microsoft SQL Server MongoDB MySQL NoSQL NumPy Object-Oriented Software Development Performance Tuning Power BI Cloud Services Tensorflow Search Technologies Amazon Simple Notification Service (SNS) SQL Databases Data Streaming Systems Integration Tableau (Software) Data Processing Google Cloud Feature Engineering Data Ingestion Pytorch Retrieval-Augmented Generation Flask (Web Framework) Large Language Models Snowflake Multi-Agent Systems Prompt Engineering Generative AI Change Data Capture Backend Fastapi Pandas Data Lakes Pyspark Scikit Learn Kubernetes Data Lineage HuggingFace Data Analytics Xgboost Apache Kafka Machine Learning Operations Virtual Agents Restful APIs Azure Synapse Analytics Looker Analytics Software Version Control Data Pipelines Docker Databricks

Job description

  • Collaborate with product owners, data scientists, data engineers, analysts, architects, and business stakeholders to identify high-value data, AI, and analytics opportunities.
  • Translate business requirements into analytical approaches, technical specifications, success metrics, acceptance criteria, and implementation plans.
  • Collect, profile, analyze, and interpret large structured, semi-structured, and unstructured datasets to uncover trends, patterns, anomalies, and business opportunities.
  • Perform exploratory data analysis, statistical analysis, hypothesis testing, correlation analysis, segmentation, and feature engineering.
  • Develop, train, tune, and evaluate machine-learning models for regression, classification, clustering, forecasting, recommendation, optimization, and anomaly detection.
  • Apply advanced statistical and machine-learning techniques, including ensemble methods, time-series modeling, causal inference, and predictive analytics.
  • Select suitable algorithms and evaluate models using precision, recall, F1 score, AUC, RMSE, MAE, explainability, fairness, and business-impact metrics.
  • Design experiments and A/B tests to validate hypotheses, measure model effectiveness, and quantify the business impact of data-driven solutions.
  • Develop business dashboards, reports, data visualizations, and self-service analytics solutions that communicate findings to technical and nontechnical stakeholders.
  • Define key performance indicators, build reusable analytical datasets, and perform ad hoc and root-cause analyses to support strategic and operational decisions.
  • Develop end-to-end Retrieval-Augmented Generation pipelines using enterprise data, embeddings, semantic search, hybrid search, reranking, vector databases, and large language models.
  • Apply prompt engineering, grounding, citation validation, structured outputs, evaluation frameworks, and guardrails to improve the accuracy and reliability of Generative AI applications.
  • Design Agentic AI solutions capable of planning, reasoning, tool selection, workflow execution, reflection, and autonomous task completion.
  • Build single-agent, multi-agent, and supervisor-agent workflows using LangGraph, LangChain, LlamaIndex, or comparable Agentic AI frameworks.
  • Integrate AI agents with enterprise APIs, databases, vector stores, documents, cloud services, analytics platforms, and business applications.
  • Implement function calling, tool calling, workflow routing, agent memory, state management, checkpointing, and Model Context Protocol integrations.
  • Develop human-in-the-loop approvals, confidence thresholds, escalation processes, fallback mechanisms, retry policies, and error-recovery controls for agentic workflows.
  • Design scalable data models and analytical schemas using MongoDB, relational databases, cloud data warehouses, data lakes.
  • Develop and maintain ETL/ELT pipelines for data ingestion, cleansing, transformation, validation, normalization, enrichment, and preparation.
  • Build and support batch, streaming, Change Data Capture, and event-driven data pipelines across operational, analytical, and AI systems.
  • Implement data-quality rules, metadata management, lineage, observability, schema-evolution handling, and governance controls.
  • Develop Python-based data and AI services and integrate machine-learning, Generative AI, and Agentic AI capabilities through APIs and reusable components.
  • Package, deploy, version, monitor, and maintain machine-learning models, AI agents, analytical applications, and data pipelines across cloud and enterprise environments.
  • Monitor model accuracy, agent behavior, data drift, model drift, pipeline health, application performance, latency, token consumption, security, and infrastructure costs.
  • Evaluate AI solutions for hallucinations, bias, prompt injection, sensitive-data exposure, unauthorized tool usage, and regulatory risks while documenting findings and coordinating delivery across cross-functional, onshore, and offshore teams.

Requirements

Data Scientist with Bachelor’s degree in Computer Science, Computer Information Systems, Information Technology, or a combination of education and experience equating to the U.S. equivalent of a Bachelor’s degree in one of the aforementioned subjects., * Strong programming skills in Python and SQL, including object-oriented programming, data structures, exception handling, debugging, and performance optimization.

  • Experience with data science libraries and machine-learning frameworks such as Pandas, NumPy, Scikit-learn, XGBoost, LightGBM, TensorFlow, and PyTorch.
  • Strong knowledge of exploratory data analysis, statistical analysis, hypothesis testing, feature engineering, causal inference, experiment design, and A/B testing.
  • Experience developing regression, classification, clustering, forecasting, recommendation, optimization, and anomaly-detection models.
  • Knowledge of model-evaluation techniques and metrics, including precision, recall, F1 score, AUC, RMSE, MAE, explainability, fairness, and business-impact measurement.
  • Experience developing dashboards, reports, and self-service analytics using Power BI, Tableau, Looker, Omni, or comparable business-intelligence platforms.
  • Hands-on knowledge of Generative AI, large language models, prompt engineering, embeddings, grounding, structured outputs, tool calling, and guardrails.
  • Experience building RAG solutions using LangChain, LangGraph, LlamaIndex, Hugging Face, or comparable document-processing and AI orchestration frameworks.
  • Knowledge of Agentic AI systems involving reasoning, planning, reflection, workflow routing, tool execution, memory, state management, and autonomous task completion.
  • Experience implementing function calling, tool calling, API integrations, checkpointing, and Model Context Protocol integrations.
  • Knowledge of vector databases and search technologies such as MongoDB Atlas Vector Search, Pinecone and Chroma.
  • Experience integrating AI models through OpenAI, Anthropic, Google Gemini, Vertex AI, AWS Bedrock, Azure OpenAI, or Hugging Face APIs.
  • Knowledge of developing Python-based backend services and REST APIs using FastAPI, Flask, or comparable frameworks.
  • Experience with data-processing and distributed-computing technologies, including Pandas, NumPy, PySpark, SQL, and data-validation frameworks.
  • Experience with relational and NoSQL databases, including PostgreSQL, MySQL, SQL Server, MongoDB, and DynamoDB.
  • Experience with cloud data warehouses, data lakes, or Lakehouse platforms such as BigQuery, Snowflake, Azure Synapse, or Databricks.
  • Experience developing ETL/ELT, batch, streaming, Change Data Capture, and event-driven data pipelines.
  • Knowledge of workflow-orchestration and messaging technologies such as Apache Airflow, Kafka, Google Cloud Pub/Sub, AWS SQS, SNS, and Event Bridge.
  • Knowledge of data-quality rules, metadata management, data lineage, observability, schema evolution, and data-governance controls.
  • Knowledge of deploying AI/ML applications using Docker, Kubernetes, and cloud platforms.
  • Understanding MLOps and LLMOps practices, including model versioning, experiment tracking, CI/CD, automated testing, controlled releases, rollback, retraining, and production monitoring.
  • Knowledge of AI security and responsible AI practices, including hallucination evaluation, bias detection, prompt-injection prevention, sensitive-data protection, explainability, model governance, privacy, and regulatory compliance.
  • Excellent analytical, problem-solving, architecture, code-review, documentation, communication, collaboration, and stakeholder-management skills.

Work location is Portland, ME with required travel to client locations throughout USA.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on ritepros.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

2:18 min

Scaling MySQL databases for massive user growth

Johannes Nicolai Johannes Nicolai +1 · LIVE

1:34 min

Bringing diverse skills to industrial data science roles

Katja Träumner

3:16 min

Terminology differences between relational and NoSQL databases

Tim Faulkes · LIVE

Videos

See all

Related articles

See all