Data scientist
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+6 more
Job description
Build risk-scoring models over synthetic tabular data, engineering features from curated medallion-layer tables. Build anomaly and outlier detection to surface irregularities in records and process data. Build optimization models for prioritization, routing, and resource allocation. Validate honestly - calibration, discrimination, stability, explainability. A correctly characterized model matters more than a flattering headline metric. Package deliverables as jobs and Asset Bundles, tracked in MLflow, and document assumptions, limitations, and what must be revalidated against real data post-ATO.
Requirements
Strong background building and deploying machine learning models. Experience with: Predictive modeling Classification Clustering Statistical analysis Feature engineering Model evaluation Experience preparing, cleaning, and curating large datasets. Hands-on experience with Databricks preferred. Experience creating synthetic datasets is a plus. Comfortable taking models from concept through production. Looking for candidates who enjoy solving business problems with data and can work independently., U.S. citizenship and active T5/SSBI federally adjudicated clearance required. Hands-on Databricks. Feature engineering on tabular and time-series data - encoding, aggregation, leakage prevention, and selection grounded in domain reasoning rather than automated search alone. Supervised learning on tabular data: gradient boosting (XGBoost/LightGBM), regularized regression, and the judgment to know when the simpler model is the right answer. Model calibration and evaluation under class imbalance - you can explain why AUC alone is insufficient for a risk score. Anomaly detection: isolation forests, autoencoders, statistical process control, or comparable - with a clear account of how you validated detections without labels. Optimization: LP/MIP or heuristic methods (OR-Tools, Pyomo, SciPy, or equivalent) applied to a real allocation or prioritization problem. Explainability (SHAP or comparable) in a decision-support context. Privacy-preserving synthetic data generation from CUI, PII, or comparably restricted source data - relational tabular data with distributional fidelity, cross-column correlations, referential integrity, and preservation of the rare-event structure that anomaly detection and risk scoring depend on. Includes an understanding of re-identification risk. Strong Python, SQL, and Spark. Government or defense contracting experience.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.clearancejobs.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?
Highest Paying Tech Companies for Developers
Making Data Warehouses Fast: A Developer’s Story
Data Engineer Salary UK