> Markdown version of [/jobs/ext/2867721-data-scientist](https://www.wearedevelopers.com/jobs/ext/2867721-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist - **Company:** Rite Pros - **Location:** Portland, ME, United States - **Contract:** Permanent contract - **Skills:** A/B Testing, Adaptable Database Systems, Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Amazon Web Services, Business Analytics Applications, Data Analysis, Application Performance Management, Automation of Tests, Microsoft Azure, BigQuery, Business Software, Cloud Computing, Cloud Database, Code Review, Computer Programming, Databases, Continuous Integration, Extract Transform Load (ETL), Data Structures, Data Visualization, Relational Databases, Software Debugging, Amazon DynamoDB, Statistical Hypothesis Testing, Python (Programming Language), PostgreSQL, Machine Learning, Enterprise Messaging Systems, Meta-Data Management, Microsoft Message Queuing, Microsoft SQL Server, MongoDB, MySQL, NoSQL, NumPy, Object-Oriented Software Development, Performance Tuning, Power BI, Cloud Services, Tensorflow, Search Technologies, Amazon Simple Notification Service (SNS), SQL Databases, Data Streaming, Systems Integration, Tableau (Software), Data Processing, Google Cloud, Feature Engineering, Data Ingestion, Pytorch, Retrieval-Augmented Generation, Flask (Web Framework), Large Language Models, Snowflake, Multi-Agent Systems, Prompt Engineering, Generative AI, Change Data Capture, Backend, Fastapi, Pandas, Data Lakes, Pyspark, Scikit Learn, Kubernetes, Data Lineage, HuggingFace, Data Analytics, Xgboost, Apache Kafka, Machine Learning Operations, Virtual Agents, Restful APIs, Azure Synapse Analytics, Looker Analytics, Software Version Control, Data Pipelines, Docker, Databricks - **Published:** September 12, 2026 - **Apply:** https://ritepros.com/data_s.php ## About the Role Data Scientist with Bachelor's degree in Computer Science, Computer Information Systems, Information Technology, or a combination of education and experience equating to the U.S. equivalent of a Bachelor's degree in one of the aforementioned subjects., * Strong programming skills in Python and SQL, including object-oriented programming, data structures, exception handling, debugging, and performance optimization. * Experience with data science libraries and machine-learning frameworks such as Pandas, NumPy, Scikit-learn, XGBoost, LightGBM, TensorFlow, and PyTorch. * Strong knowledge of exploratory data analysis, statistical analysis, hypothesis testing, feature engineering, causal inference, experiment design, and A/B testing. * Experience developing regression, classification, clustering, forecasting, recommendation, optimization, and anomaly-detection models. * Knowledge of model-evaluation techniques and metrics, including precision, recall, F1 score, AUC, RMSE, MAE, explainability, fairness, and business-impact measurement. * Experience developing dashboards, reports, and self-service analytics using Power BI, Tableau, Looker, Omni, or comparable business-intelligence platforms. * Hands-on knowledge of Generative AI, large language models, prompt engineering, embeddings, grounding, structured outputs, tool calling, and guardrails. * Experience building RAG solutions using LangChain, LangGraph, LlamaIndex, Hugging Face, or comparable document-processing and AI orchestration frameworks. * Knowledge of Agentic AI systems involving reasoning, planning, reflection, workflow routing, tool execution, memory, state management, and autonomous task completion. * Experience implementing function calling, tool calling, API integrations, checkpointing, and Model Context Protocol integrations. * Knowledge of vector databases and search technologies such as MongoDB Atlas Vector Search, Pinecone and Chroma. * Experience integrating AI models through OpenAI, Anthropic, Google Gemini, Vertex AI, AWS Bedrock, Azure OpenAI, or Hugging Face APIs. * Knowledge of developing Python-based backend services and REST APIs using FastAPI, Flask, or comparable frameworks. * Experience with data-processing and distributed-computing technologies, including Pandas, NumPy, PySpark, SQL, and data-validation frameworks. * Experience with relational and NoSQL databases, including PostgreSQL, MySQL, SQL Server, MongoDB, and DynamoDB. * Experience with cloud data warehouses, data lakes, or Lakehouse platforms such as BigQuery, Snowflake, Azure Synapse, or Databricks. * Experience developing ETL/ELT, batch, streaming, Change Data Capture, and event-driven data pipelines. * Knowledge of workflow-orchestration and messaging technologies such as Apache Airflow, Kafka, Google Cloud Pub/Sub, AWS SQS, SNS, and Event Bridge. * Knowledge of data-quality rules, metadata management, data lineage, observability, schema evolution, and data-governance controls. * Knowledge of deploying AI/ML applications using Docker, Kubernetes, and cloud platforms. * Understanding MLOps and LLMOps practices, including model versioning, experiment tracking, CI/CD, automated testing, controlled releases, rollback, retraining, and production monitoring. * Knowledge of AI security and responsible AI practices, including hallucination evaluation, bias detection, prompt-injection prevention, sensitive-data protection, explainability, model governance, privacy, and regulatory compliance. * Excellent analytical, problem-solving, architecture, code-review, documentation, communication, collaboration, and stakeholder-management skills. Work location is Portland, ME with required travel to client locations throughout USA. ## Description * Collaborate with product owners, data scientists, data engineers, analysts, architects, and business stakeholders to identify high-value data, AI, and analytics opportunities. * Translate business requirements into analytical approaches, technical specifications, success metrics, acceptance criteria, and implementation plans. * Collect, profile, analyze, and interpret large structured, semi-structured, and unstructured datasets to uncover trends, patterns, anomalies, and business opportunities. * Perform exploratory data analysis, statistical analysis, hypothesis testing, correlation analysis, segmentation, and feature engineering. * Develop, train, tune, and evaluate machine-learning models for regression, classification, clustering, forecasting, recommendation, optimization, and anomaly detection. * Apply advanced statistical and machine-learning techniques, including ensemble methods, time-series modeling, causal inference, and predictive analytics. * Select suitable algorithms and evaluate models using precision, recall, F1 score, AUC, RMSE, MAE, explainability, fairness, and business-impact metrics. * Design experiments and A/B tests to validate hypotheses, measure model effectiveness, and quantify the business impact of data-driven solutions. * Develop business dashboards, reports, data visualizations, and self-service analytics solutions that communicate findings to technical and nontechnical stakeholders. * Define key performance indicators, build reusable analytical datasets, and perform ad hoc and root-cause analyses to support strategic and operational decisions. * Develop end-to-end Retrieval-Augmented Generation pipelines using enterprise data, embeddings, semantic search, hybrid search, reranking, vector databases, and large language models. * Apply prompt engineering, grounding, citation validation, structured outputs, evaluation frameworks, and guardrails to improve the accuracy and reliability of Generative AI applications. * Design Agentic AI solutions capable of planning, reasoning, tool selection, workflow execution, reflection, and autonomous task completion. * Build single-agent, multi-agent, and supervisor-agent workflows using LangGraph, LangChain, LlamaIndex, or comparable Agentic AI frameworks. * Integrate AI agents with enterprise APIs, databases, vector stores, documents, cloud services, analytics platforms, and business applications. * Implement function calling, tool calling, workflow routing, agent memory, state management, checkpointing, and Model Context Protocol integrations. * Develop human-in-the-loop approvals, confidence thresholds, escalation processes, fallback mechanisms, retry policies, and error-recovery controls for agentic workflows. * Design scalable data models and analytical schemas using MongoDB, relational databases, cloud data warehouses, data lakes. * Develop and maintain ETL/ELT pipelines for data ingestion, cleansing, transformation, validation, normalization, enrichment, and preparation. * Build and support batch, streaming, Change Data Capture, and event-driven data pipelines across operational, analytical, and AI systems. * Implement data-quality rules, metadata management, lineage, observability, schema-evolution handling, and governance controls. * Develop Python-based data and AI services and integrate machine-learning, Generative AI, and Agentic AI capabilities through APIs and reusable components. * Package, deploy, version, monitor, and maintain machine-learning models, AI agents, analytical applications, and data pipelines across cloud and enterprise environments. * Monitor model accuracy, agent behavior, data drift, model drift, pipeline health, application performance, latency, token consumption, security, and infrastructure costs. * Evaluate AI solutions for hallucinations, bias, prompt injection, sensitive-data exposure, unauthorized tool usage, and regulatory risks while documenting findings and coordinating delivery across cross-functional, onshore, and offshore teams. ## Related Videos - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [MySQL Protocol Features You Should Be Aware Of](https://www.wearedevelopers.com/videos/100267-mysql-protocol-features-you-should-be-aware-of) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [NoSQL Data Modeling for Front-end Developers](https://www.wearedevelopers.com/videos/297-nosql-data-modeling-for-front-end-developers) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)