Data Scientist

MarineTraffic
United States
3 days ago
Apply on www2.jobdiva.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$26,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Encodings Data Integrity Python (Programming Language) SciPy Statistical Process Control (SPC) SQL Databases Supervised Learning Data Processing Feature Engineering Large Language Models SHAP (Shapley Additive Explanations)
+4 more
Apache Spark Xgboost Machine Learning Operations Databricks

Job description

Build risk-scoring models over synthetic tabular data, engineering features from curated medallion-layer tables. Build anomaly and outlier detection to surface irregularities in records and process data. Build optimization models for prioritization, routing, and resource allocation. Validate honestly - calibration, discrimination, stability, explainability. A correctly characterized model matters more than a flattering headline metric. Package deliverables as jobs and Asset Bundles, tracked in MLflow, and document assumptions, limitations, and what must be revalidated against real data post-ATO.

Requirements

U.S. citizenship and active T5/SSBI federally adjudicated clearance required. Hands-on Databricks. Feature engineering on tabular and time-series data - encoding, aggregation, leakage prevention, and selection grounded in domain reasoning rather than automated search alone. Supervised learning on tabular data: gradient boosting (XGBoost/LightGBM), regularized regression, and the judgment to know when the simpler model is the right answer. Model calibration and evaluation under class imbalance - you can explain why AUC alone is insufficient for a risk score. Anomaly detection: isolation forests, autoencoders, statistical process control, or comparable - with a clear account of how you validated detections without labels. Optimization: LP/MIP or heuristic methods (OR-Tools, Pyomo, SciPy, or equivalent) applied to a real allocation or prioritization problem. Explainability (SHAP or comparable) in a decision-support context. Privacy-preserving synthetic data generation from CUI, PII, or comparably restricted source data - relational tabular data with distributional fidelity, cross-column correlations, referential integrity, and preservation of the rare-event structure that anomaly detection and risk scoring depend on. Includes an understanding of re-identification risk. Strong Python, SQL, and Spark. Government or defense contracting experience., Modeling on federal investigative, vetting, fraud, or insider-threat data. Direct experience with FedRAMP, NIST 800-171, CMMC L2, or CUI handling. Familiarity with LLM/GenAI workflows - useful for collaboration with a peer document-intelligence workstream, but secondary to the core ML skill set. H2O (Driverless AI, H2O-3). MLflow, Databricks Asset Bundles, Unity Catalog. Fairness / adverse-impact analysis in a regulated or decision-support setting.

Soft Skills Self-directed execution against a fixed milestone with minimal oversight. Honest reporting of model behavior - comfortable stating what synthetic-data performance does and does not establish about real-world accuracy. Collaboration across technical and non-technical teams. Clear documentation and active knowledge transfer.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www2.jobdiva.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

4:36 min

Evaluating model performance and utilizing self-supervised learning

Humera Minhas +1 · World Congress 2022

4:54 min

Development history of scientific computation libraries and PyViz tools

Radovan Kavický · LIVE

3:17 min

Optimizing character encoding with Kim variable byte encoding

Douglas Crockford Douglas Crockford · World Congress 2024

1:34 min

Bringing diverse skills to industrial data science roles

Katja Träumner

1:09 min

Configuring synthetic data for safe interactive programming

Mingshen Sun Mingshen Sun · World Congress 2024

Videos

See all

Related articles

See all