AI/ML Data Engineer

PRECISE SOFTWARE SOLUTIONS INC.
Washington, DC, United States
1 day ago
Apply on diversityjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
1 year minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Airflow Amazon Web Services Microsoft Azure Encodings Information Systems Data Dictionary Information Engineering Data Governance Extract Transform Load (ETL) Data Migration Data Security
+26 more
Database Development Federal Information Processing Standards (FIPS) Python (Programming Language) Machine Learning Meta-Data Management SQL Databases Management of Software Versions Feature Store Google Cloud Cloud Platform System Normalized Discounted Cumulative Gain Retrieval-Augmented Generation Large Language Models Snowflake Apache Spark Data Poisoning Generative AI Git Pandas Pgvector Information Technology Integration Frameworks Machine Learning Operations Data Pipelines OpenSearch Databricks

Job description

  • Design and maintain data ingestion, transformation, and processing pipelines (ETL/ELT) for AI training, evaluation, retrieval, and operations, including support for data migration and cleansing.
  • Curate, validate, and version datasets, and maintain dataset inventories, metadata, lineage, provenance, and ingestion logs.
  • Implement automated data-quality checks for duplication, schema changes, completeness, and freshness, and maintain dataset quality scorecards and drift reports.
  • Design data models, vector stores, and embedding schemas for Retrieval-Augmented Generation (RAG) knowledge bases, and re-index content when sources change.
  • Measure retrieval and model quality against established baselines using metrics such as precision/recall, MRR, NDCG, and context relevance.
  • Prepare data-related deliverables, including AI model cards, ML and AI pipeline documentation, RAG/AI Pipeline Evaluation Reports, data dictionaries and embedding schema documentation, and responses to Government data calls.
  • Build secure structured-data access for AI applications (e.g., natural-language-to-SQL with query validation and role-based authorization), and support dashboards and operational analytics.
  • Monitor data and ML pipelines, troubleshoot failures, and support root-cause analysis, while keeping all data in FedRAMP-authorized cloud regions with FIPS-validated encryption.
  • Support Responsible AI practices by preparing representative evaluation datasets, testing AI outputs for bias, accuracy, and hallucination, and documenting results to meet federal AI governance requirements.
  • Secure the AI data path, from source datasets and embeddings to prompts and logs, and support AI risk testing such as data poisoning.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, Mathematics, Information Systems, or a related field and 5+ years of relevant experience. Equivalent experience may substitute for the degree.
  • 5+ years of hands-on data engineering or database development experience, including data modeling, SQL, and ETL/ELT pipeline development.
  • Strong proficiency in Python and experience with data processing frameworks (e.g., pandas, Spark) and workflow orchestration tools (e.g., Airflow).
  • 1+ year of experience preparing data for Generative AI or machine learning, such as embeddings, vector databases, RAG knowledge bases, or training and evaluation datasets.
  • Experience implementing data quality, lineage, metadata management, and data governance controls.
  • Experience protecting sensitive data (PII), including masking, minimization, and access controls.
  • Experience with a major cloud data platform (AWS, Azure, or Google Cloud), Git, and CI/CD tools.
  • Must be a U.S. citizen or lawful permanent resident (green card holder).
  • Must reside in the Washington, DC metropolitan area and be able to work on-site at Government offices.
  • Must be able to obtain and maintain a Public Trust background investigation., * Master’s degree in a related field and 7+ years of data engineering experience, including support of federal agency programs.
  • Experience evaluating retrieval quality and building RAG pipelines with vector stores (e.g., pgvector, OpenSearch).
  • Experience with MLOps tooling, including model and dataset versioning, feature stores, or model registries.
  • Generative AI, LLM security, or data certification (e.g., Databricks Generative AI Engineer, Snowflake SnowPro, AWS, Microsoft Azure, or Google Cloud).
  • An active Public Trust or prior federal background investigation.

Benefits & conditions

  • Comprehensive Health Benefits (Medical, Dental and Vision)
  • Flexible Spending Accounts (FSA) & Health Savings Account (HSA)
  • Retirement Plan with 4% match and discretionary match at year end
  • Paid Time Off (PTO): 15 days of PTO accrued per year; 7 holidays+ 3 Floating holidays; 2 Innovation days (paid training days)
  • Short Term and Long-Term Disability
  • Paid Parental Leave
  • Paid Jury Duty leave
  • Life and AD&D Insurance
  • Critical Illness Insurance
  • Training and Development
  • Wellness Incentives & Discount programs
  • Employee Referral Program
  • Annual Charity Donation Match
  • Awards and Recognition

About the company

Precise Software Solutions, Inc. is a mission-focused technology services company delivering secure digital platforms, infrastructure, and operational IT services to government organizations. A CMMI Level 3-appraised company, Precise partners with agency technology leaders and solution providers to design, build, operate, and modernize enterprise IT solutions that support critical public missions combining agility, innovation, and performance to deliver measurable results.

Precise specializes in cloud and hybrid infrastructure, platform engineering, security operations and compliance, application modernization, and data platforms and analytics. The company is known for its agile, delivery-driven approach and innovative engineering practices, applying operational rigor and performance-focused execution to improve system resilience, security, and scalability across complex government environments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on diversityjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

2:19 min

Introduction to Apache Airflow for advanced orchestration

Alan Mazankiewicz · LIVE

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all