Data Scientist
Rite Pros
Portland, ME, United States
8 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on ritepros.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source
Tech stack
A/B Testing
Adaptable Database Systems
Application Programming Interfaces (APIs)
Artificial Intelligence
Airflow
Amazon Web Services
Business Analytics Applications
Data Analysis
Application Performance Management
Automation of Tests
Microsoft Azure
BigQuery
+70 more
Business Software
Cloud Computing
Cloud Database
Code Review
Computer Programming
Databases
Continuous Integration
Extract Transform Load (ETL)
Data Structures
Data Visualization
Relational Databases
Software Debugging
Amazon DynamoDB
Statistical Hypothesis Testing
Python (Programming Language)
PostgreSQL
Machine Learning
Enterprise Messaging Systems
Meta-Data Management
Microsoft Message Queuing
Microsoft SQL Server
MongoDB
MySQL
NoSQL
NumPy
Object-Oriented Software Development
Performance Tuning
Power BI
Cloud Services
Tensorflow
Search Technologies
Amazon Simple Notification Service (SNS)
SQL Databases
Data Streaming
Systems Integration
Tableau (Software)
Data Processing
Google Cloud
Feature Engineering
Data Ingestion
Pytorch
Retrieval-Augmented Generation
Flask (Web Framework)
Large Language Models
Snowflake
Multi-Agent Systems
Prompt Engineering
Generative AI
Change Data Capture
Backend
Fastapi
Pandas
Data Lakes
Pyspark
Scikit Learn
Kubernetes
Data Lineage
HuggingFace
Data Analytics
Xgboost
Apache Kafka
Machine Learning Operations
Virtual Agents
Restful APIs
Azure Synapse Analytics
Looker Analytics
Software Version Control
Data Pipelines
Docker
Databricks
Job description
- Collaborate with product owners, data scientists, data engineers, analysts, architects, and business stakeholders to identify high-value data, AI, and analytics opportunities.
- Translate business requirements into analytical approaches, technical specifications, success metrics, acceptance criteria, and implementation plans.
- Collect, profile, analyze, and interpret large structured, semi-structured, and unstructured datasets to uncover trends, patterns, anomalies, and business opportunities.
- Perform exploratory data analysis, statistical analysis, hypothesis testing, correlation analysis, segmentation, and feature engineering.
- Develop, train, tune, and evaluate machine-learning models for regression, classification, clustering, forecasting, recommendation, optimization, and anomaly detection.
- Apply advanced statistical and machine-learning techniques, including ensemble methods, time-series modeling, causal inference, and predictive analytics.
- Select suitable algorithms and evaluate models using precision, recall, F1 score, AUC, RMSE, MAE, explainability, fairness, and business-impact metrics.
- Design experiments and A/B tests to validate hypotheses, measure model effectiveness, and quantify the business impact of data-driven solutions.
- Develop business dashboards, reports, data visualizations, and self-service analytics solutions that communicate findings to technical and nontechnical stakeholders.
- Define key performance indicators, build reusable analytical datasets, and perform ad hoc and root-cause analyses to support strategic and operational decisions.
- Develop end-to-end Retrieval-Augmented Generation pipelines using enterprise data, embeddings, semantic search, hybrid search, reranking, vector databases, and large language models.
- Apply prompt engineering, grounding, citation validation, structured outputs, evaluation frameworks, and guardrails to improve the accuracy and reliability of Generative AI applications.
- Design Agentic AI solutions capable of planning, reasoning, tool selection, workflow execution, reflection, and autonomous task completion.
- Build single-agent, multi-agent, and supervisor-agent workflows using LangGraph, LangChain, LlamaIndex, or comparable Agentic AI frameworks.
- Integrate AI agents with enterprise APIs, databases, vector stores, documents, cloud services, analytics platforms, and business applications.
- Implement function calling, tool calling, workflow routing, agent memory, state management, checkpointing, and Model Context Protocol integrations.
- Develop human-in-the-loop approvals, confidence thresholds, escalation processes, fallback mechanisms, retry policies, and error-recovery controls for agentic workflows.
- Design scalable data models and analytical schemas using MongoDB, relational databases, cloud data warehouses, data lakes.
- Develop and maintain ETL/ELT pipelines for data ingestion, cleansing, transformation, validation, normalization, enrichment, and preparation.
- Build and support batch, streaming, Change Data Capture, and event-driven data pipelines across operational, analytical, and AI systems.
- Implement data-quality rules, metadata management, lineage, observability, schema-evolution handling, and governance controls.
- Develop Python-based data and AI services and integrate machine-learning, Generative AI, and Agentic AI capabilities through APIs and reusable components.
- Package, deploy, version, monitor, and maintain machine-learning models, AI agents, analytical applications, and data pipelines across cloud and enterprise environments.
- Monitor model accuracy, agent behavior, data drift, model drift, pipeline health, application performance, latency, token consumption, security, and infrastructure costs.
- Evaluate AI solutions for hallucinations, bias, prompt injection, sensitive-data exposure, unauthorized tool usage, and regulatory risks while documenting findings and coordinating delivery across cross-functional, onshore, and offshore teams.
Requirements
Data Scientist with Bachelor’s degree in Computer Science, Computer Information Systems, Information Technology, or a combination of education and experience equating to the U.S. equivalent of a Bachelor’s degree in one of the aforementioned subjects., * Strong programming skills in Python and SQL, including object-oriented programming, data structures, exception handling, debugging, and performance optimization.
- Experience with data science libraries and machine-learning frameworks such as Pandas, NumPy, Scikit-learn, XGBoost, LightGBM, TensorFlow, and PyTorch.
- Strong knowledge of exploratory data analysis, statistical analysis, hypothesis testing, feature engineering, causal inference, experiment design, and A/B testing.
- Experience developing regression, classification, clustering, forecasting, recommendation, optimization, and anomaly-detection models.
- Knowledge of model-evaluation techniques and metrics, including precision, recall, F1 score, AUC, RMSE, MAE, explainability, fairness, and business-impact measurement.
- Experience developing dashboards, reports, and self-service analytics using Power BI, Tableau, Looker, Omni, or comparable business-intelligence platforms.
- Hands-on knowledge of Generative AI, large language models, prompt engineering, embeddings, grounding, structured outputs, tool calling, and guardrails.
- Experience building RAG solutions using LangChain, LangGraph, LlamaIndex, Hugging Face, or comparable document-processing and AI orchestration frameworks.
- Knowledge of Agentic AI systems involving reasoning, planning, reflection, workflow routing, tool execution, memory, state management, and autonomous task completion.
- Experience implementing function calling, tool calling, API integrations, checkpointing, and Model Context Protocol integrations.
- Knowledge of vector databases and search technologies such as MongoDB Atlas Vector Search, Pinecone and Chroma.
- Experience integrating AI models through OpenAI, Anthropic, Google Gemini, Vertex AI, AWS Bedrock, Azure OpenAI, or Hugging Face APIs.
- Knowledge of developing Python-based backend services and REST APIs using FastAPI, Flask, or comparable frameworks.
- Experience with data-processing and distributed-computing technologies, including Pandas, NumPy, PySpark, SQL, and data-validation frameworks.
- Experience with relational and NoSQL databases, including PostgreSQL, MySQL, SQL Server, MongoDB, and DynamoDB.
- Experience with cloud data warehouses, data lakes, or Lakehouse platforms such as BigQuery, Snowflake, Azure Synapse, or Databricks.
- Experience developing ETL/ELT, batch, streaming, Change Data Capture, and event-driven data pipelines.
- Knowledge of workflow-orchestration and messaging technologies such as Apache Airflow, Kafka, Google Cloud Pub/Sub, AWS SQS, SNS, and Event Bridge.
- Knowledge of data-quality rules, metadata management, data lineage, observability, schema evolution, and data-governance controls.
- Knowledge of deploying AI/ML applications using Docker, Kubernetes, and cloud platforms.
- Understanding MLOps and LLMOps practices, including model versioning, experiment tracking, CI/CD, automated testing, controlled releases, rollback, retraining, and production monitoring.
- Knowledge of AI security and responsible AI practices, including hallucination evaluation, bias detection, prompt-injection prevention, sensitive-data protection, explainability, model governance, privacy, and regulatory compliance.
- Excellent analytical, problem-solving, architecture, code-review, documentation, communication, collaboration, and stakeholder-management skills.
Work location is Portland, ME with required travel to client locations throughout USA.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on ritepros.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
EM
Eli McGarvie
over 3 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
LM
Luis Minvielle
How to Become an AI Engineer
almost 3 years ago
EM
Eli McGarvie
Data Analyst Salary in the UK
about 3 years ago
DS
Dhannush Subramani
Top Big Data Technologies That You Need to Know
about 4 years ago
CH
Chris Heilmann
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
almost 2 years ago