> Markdown version of [/jobs/ext/194378-data-scientist](https://www.wearedevelopers.com/jobs/ext/194378-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist - **Company:** W. H. GREEN & SONS, INC. - **Location:** Portland, ME, United States - **Contract:** Permanent contract - **Skills:** Clean Code Principles, Amazon Web Services, Amazon S3, Data Analysis, Big Data, BigQuery, Cloud Storage, Data Visualization, Relational Databases, DevOps, Programming Tools, Distributed Computing Environment, Distributed Data Store, Statistical Hypothesis Testing, Python (Programming Language), Machine Learning, MongoDB, Regression Analysis, NLTK (NLP Analysis), NumPy, OpenCV, Tensorflow, Standard Sql, Software Engineering, SQL Databases, Jupyter Notebook, Google Cloud, Cloud Platform System, Feature Engineering, Pytorch, Git, Pandas, Matplotlib, Scikit Learn, Information Technology, HuggingFace, Plotly, Bitbucket, Gensim, Spacy, Software Version Control - **Published:** May 30, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=53938f65e1e5f649 ## About the Role Do you have experience in spaCy?, Do you have a Bachelor's degree?, Data Scientist with Bachelor's Degree in Computer Science, Computer Information Systems, Information Technology, or a combination of education and experience equating to the U.S. equivalent of a Bachelor's degree in one of the aforementioned subjects. ## Description * Analyze large-scale structured and semi-structured datasets to uncover patterns, trends, and insights that support business and product decisions. * Develop, train, and evaluate machine learning models for use cases such as prediction, classification, anomaly detection, and forecasting. * Perform exploratory data analysis (EDA) to understand data distributions, detect anomalies, and guide feature engineering strategies. * Apply statistical techniques including hypothesis testing, regression analysis, and probability modeling to validate results and support decision-making. * Design and implement feature engineering pipelines to transform raw data into meaningful inputs for machine learning models. * Build and compare multiple models using appropriate evaluation metrics (accuracy, precision, recall, F1-score, ROC-AUC) and optimize performance through tuning. * Work with large datasets using distributed computing frameworks or cloud-based platforms to ensure scalability and efficiency. * Develop data visualizations, dashboards, and reports to effectively communicate analytical findings to technical and non-technical stakeholders. * Collaborate with cross-functional teams including product managers, engineers, and business teams to translate business problems into data-driven solutions. * Support the deployment of machine learning models into production by working with engineering teams and ensuring models meet performance and reliability standards. * Monitor model performance over time and assist in updating models based on new data and changing business requirements. * Write clean, modular, and maintainable code following best practices in software development and version control. * Document analytical workflows, model assumptions, and results to ensure reproducibility and knowledge sharing across teams. Technologies / Environment involved: * Distributed storage: AWS Cloud Storage (S3), Google Cloud (GCP - Cloud Storage, BigQuery) * Database management: MongoDB, SQL (Relational Databases) * Machine learning: TensorFlow, PyTorch, Scikit-learn, NumPy, Pandas; exposure to SpaCy, NLTK, HuggingFace, Gensim, OpenCV * Programming Languages:Python, SQL * Data Visualization:Matplotlib, Seaborn, Plotly * Development Tools: Jupyter Notebook / JupyterLab * DevOps Tools:Git, Bitbucket ## Related Videos - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Vectorize all the things! Using linear algebra and NumPy to make your Python code lightning fast.](https://www.wearedevelopers.com/videos/562-vectorize-all-the-things-using-linear-algebra-and-numpy-to-make-your-python-code-lightning-fast) - [Python Data Visualization @ Deepnote (w/ PyViz overview)](https://www.wearedevelopers.com/videos/113-python-data-visualization-deepnote-w-pyviz-overview) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [How to implement convenient Python bindings to C++](https://www.wearedevelopers.com/videos/618-how-to-implement-convenient-python-bindings-to-c) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [The 13 Best Python Libraries for Developers in 2025](https://www.wearedevelopers.com/magazine/371-the-13-best-python-libraries-for-developers-in-2025) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Best Coding Boot Camps in Germany](https://www.wearedevelopers.com/magazine/237-best-coding-boot-camps-in-germany) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story)