> Markdown version of [/jobs/ext/3555112-data-scientist](https://www.wearedevelopers.com/jobs/ext/3555112-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist - **Company:** MarineTraffic - **Location:** United States - **Salary:** $26,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Encodings, Data Integrity, Python (Programming Language), SciPy, Statistical Process Control (SPC), SQL Databases, Supervised Learning, Data Processing, Feature Engineering, Large Language Models, SHAP (Shapley Additive Explanations), Apache Spark, Xgboost, Machine Learning Operations, Databricks - **Published:** October 2, 2026 - **Apply:** https://www2.jobdiva.com/portal/?a=j7jdnwgck4wulo816y7l865pzq556q05f5cqdk65b0o0dpy0m5p196w9a9ln1g3w&compid=0/jobs/33188362#/jobs/33188362 ## About the Role U.S. citizenship and active T5/SSBI federally adjudicated clearance required. Hands-on Databricks. Feature engineering on tabular and time-series data - encoding, aggregation, leakage prevention, and selection grounded in domain reasoning rather than automated search alone. Supervised learning on tabular data: gradient boosting (XGBoost/LightGBM), regularized regression, and the judgment to know when the simpler model is the right answer. Model calibration and evaluation under class imbalance - you can explain why AUC alone is insufficient for a risk score. Anomaly detection: isolation forests, autoencoders, statistical process control, or comparable - with a clear account of how you validated detections without labels. Optimization: LP/MIP or heuristic methods (OR-Tools, Pyomo, SciPy, or equivalent) applied to a real allocation or prioritization problem. Explainability (SHAP or comparable) in a decision-support context. Privacy-preserving synthetic data generation from CUI, PII, or comparably restricted source data - relational tabular data with distributional fidelity, cross-column correlations, referential integrity, and preservation of the rare-event structure that anomaly detection and risk scoring depend on. Includes an understanding of re-identification risk. Strong Python, SQL, and Spark. Government or defense contracting experience., Modeling on federal investigative, vetting, fraud, or insider-threat data. Direct experience with FedRAMP, NIST 800-171, CMMC L2, or CUI handling. Familiarity with LLM/GenAI workflows - useful for collaboration with a peer document-intelligence workstream, but secondary to the core ML skill set. H2O (Driverless AI, H2O-3). MLflow, Databricks Asset Bundles, Unity Catalog. Fairness / adverse-impact analysis in a regulated or decision-support setting. Soft Skills Self-directed execution against a fixed milestone with minimal oversight. Honest reporting of model behavior - comfortable stating what synthetic-data performance does and does not establish about real-world accuracy. Collaboration across technical and non-technical teams. Clear documentation and active knowledge transfer. ## Description Build risk-scoring models over synthetic tabular data, engineering features from curated medallion-layer tables. Build anomaly and outlier detection to surface irregularities in records and process data. Build optimization models for prioritization, routing, and resource allocation. Validate honestly - calibration, discrimination, stability, explainability. A correctly characterized model matters more than a flattering headline metric. Package deliverables as jobs and Asset Bundles, tracked in MLflow, and document assumptions, limitations, and what must be revalidated against real data post-ATO. ## Related Videos - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Shoot for the moon - machine learning for automated online ad detection](https://www.wearedevelopers.com/videos/502-shoot-for-the-moon-machine-learning-for-automated-online-ad-detection) - [Python Data Visualization @ Deepnote (w/ PyViz overview)](https://www.wearedevelopers.com/videos/113-python-data-visualization-deepnote-w-pyviz-overview) - [A Brief History of Data Storage](https://www.wearedevelopers.com/videos/974-a-brief-history-of-data-storage) - [The pitfalls of Deep Learning - When Neural Networks are not the solution](https://www.wearedevelopers.com/videos/14-the-pitfalls-of-deep-learning-when-neural-networks-are-not-the-solution) - [JSON and Beyond](https://www.wearedevelopers.com/videos/968-json-and-beyond) ## Related Articles - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [ I Gave a Video Editor More Autonomy Than a Trading Bot. On Purpose.](https://www.wearedevelopers.com/magazine/773-i-gave-a-video-editor-more-autonomy-than-a-trading-bot-on-purpose) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know)