Data Scientist III
Role details
Job location
Tech stack
Job description
Our client is hiring a contract Data Scientist to support critical missions within a high-trust federal environment. This role focuses on building and validating predictive and prescriptive models on a greenfield data and AI platform. The work is hands-on and designed for transition-fully documented, portable, and transferred to the internal team. This is a 12+ month contract with the possibility of extension. An Active TS Clearance is required.
- Build risk-scoring models over synthetic tabular data, engineering features from curated medallion-layer tables.
- Build anomaly and outlier detection models to surface irregularities in records and process data.
- Build optimization models for prioritization, routing, and resource allocation.
- Validate models through calibration, discrimination, stability, and explainability. A correctly characterized model matters more than a flattering headline metric.
- Package deliverables as jobs and Asset Bundles, tracked in MLflow, and document assumptions, limitations, and what must be revalidated against real data post-ATO.
Requirements
-
Hands-on experience with Databricks.
-
Experience with feature engineering on tabular and time-series data, including encoding, aggregation, leakage prevention, and selection grounded in domain reasoning rather than automated search alone.
-
Strong experience with supervised learning on tabular data, including gradient boosting (XGBoost/LightGBM), regularized regression, and the judgment to determine when a simpler model is appropriate.
-
Experience with model calibration and evaluation under class imbalance, with the ability to explain why AUC alone is insufficient for a risk score.
-
Experience with anomaly detection using isolation forests, autoencoders, statistical process control, or comparable techniques, along with validation approaches when labeled data is unavailable.
-
Experience with optimization techniques including LP/MIP or heuristic methods (OR-Tools, Pyomo, SciPy, or equivalent) applied to real-world allocation or prioritization problems.
-
Experience implementing explainability techniques such as SHAP or comparable methods in decision-support environments.
-
Experience with privacy-preserving synthetic data generation from CUI, PII, or similarly restricted source data, including relational tabular data with distributional fidelity, cross-column correlations, referential integrity, preservation of rare-event structures, and an understanding of re-identification risk.
-
Strong programming skills in Python, SQL, and Spark.
-
Government or defense contracting experience.
-
Experience modeling federal investigative, vetting, fraud, or insider-threat data.
-
Experience with FedRAMP, NIST 800-171, CMMC Level 2, or CUI handling.
-
Familiarity with LLM/GenAI workflows to support collaboration with document intelligence initiatives.
-
Experience with H2O (Driverless AI, H2O-3).
-
Experience with MLflow, Databricks Asset Bundles, and Unity Catalog.
-
Experience performing fairness and adverse-impact analysis in regulated or decision-support environments.
-
Self-directed execution against fixed milestones with minimal supervision.
-
Honest reporting of model behavior, including clear communication of what synthetic-data performance does and does not establish about real-world accuracy.
-
Strong collaboration skills across technical and non-technical teams.
-
Excellent documentation and knowledge transfer abilities.