Data scientist

Nastech Global, Inc.
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English

Job location

Remote

Tech stack

Big Data
Cluster Analysis
Encodings
Data Integrity
Python
Machine Learning
SciPy
Statistical Process Control (SPC)
SQL Databases
Supervised Learning
Data Processing
Feature Engineering
Spark
Model Validation
XGBoost
Machine Learning Operations
Databricks
Data Generation

Job description

Build risk-scoring models over synthetic tabular data, engineering features from curated medallion-layer tables. Build anomaly and outlier detection to surface irregularities in records and process data. Build optimization models for prioritization, routing, and resource allocation. Validate honestly - calibration, discrimination, stability, explainability. A correctly characterized model matters more than a flattering headline metric. Package deliverables as jobs and Asset Bundles, tracked in MLflow, and document assumptions, limitations, and what must be revalidated against real data post-ATO.

Requirements

Strong background building and deploying machine learning models. Experience with: Predictive modeling Classification Clustering Statistical analysis Feature engineering Model evaluation Experience preparing, cleaning, and curating large datasets. Hands-on experience with Databricks preferred. Experience creating synthetic datasets is a plus. Comfortable taking models from concept through production. Looking for candidates who enjoy solving business problems with data and can work independently., U.S. citizenship and active T5/SSBI federally adjudicated clearance required. Hands-on Databricks. Feature engineering on tabular and time-series data - encoding, aggregation, leakage prevention, and selection grounded in domain reasoning rather than automated search alone. Supervised learning on tabular data: gradient boosting (XGBoost/LightGBM), regularized regression, and the judgment to know when the simpler model is the right answer. Model calibration and evaluation under class imbalance - you can explain why AUC alone is insufficient for a risk score. Anomaly detection: isolation forests, autoencoders, statistical process control, or comparable - with a clear account of how you validated detections without labels. Optimization: LP/MIP or heuristic methods (OR-Tools, Pyomo, SciPy, or equivalent) applied to a real allocation or prioritization problem. Explainability (SHAP or comparable) in a decision-support context. Privacy-preserving synthetic data generation from CUI, PII, or comparably restricted source data - relational tabular data with distributional fidelity, cross-column correlations, referential integrity, and preservation of the rare-event structure that anomaly detection and risk scoring depend on. Includes an understanding of re-identification risk. Strong Python, SQL, and Spark. Government or defense contracting experience.

Apply for this position