> Markdown version of [/jobs/ext/3538115-ai-ml-data-engineer](https://www.wearedevelopers.com/jobs/ext/3538115-ai-ml-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI/ML Data Engineer - **Company:** PRECISE SOFTWARE SOLUTIONS INC. - **Location:** Washington, DC, United States - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Microsoft Azure, Encodings, Information Systems, Data Dictionary, Information Engineering, Data Governance, Extract Transform Load (ETL), Data Migration, Data Security, Database Development, Federal Information Processing Standards (FIPS), Python (Programming Language), Machine Learning, Meta-Data Management, SQL Databases, Management of Software Versions, Feature Store, Google Cloud, Cloud Platform System, Normalized Discounted Cumulative Gain, Retrieval-Augmented Generation, Large Language Models, Snowflake, Apache Spark, Data Poisoning, Generative AI, Git, Pandas, Pgvector, Information Technology, Integration Frameworks, Machine Learning Operations, Data Pipelines, OpenSearch, Databricks - **Published:** October 2, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/18496286?backUrl=%2Fcareer%2F18496286%2FAi-Ml-Data-Engineer-D-C-Washington ## About the Role * Bachelor's degree in Computer Science, Engineering, Mathematics, Information Systems, or a related field and 5+ years of relevant experience. Equivalent experience may substitute for the degree. * 5+ years of hands-on data engineering or database development experience, including data modeling, SQL, and ETL/ELT pipeline development. * Strong proficiency in Python and experience with data processing frameworks (e.g., pandas, Spark) and workflow orchestration tools (e.g., Airflow). * 1+ year of experience preparing data for Generative AI or machine learning, such as embeddings, vector databases, RAG knowledge bases, or training and evaluation datasets. * Experience implementing data quality, lineage, metadata management, and data governance controls. * Experience protecting sensitive data (PII), including masking, minimization, and access controls. * Experience with a major cloud data platform (AWS, Azure, or Google Cloud), Git, and CI/CD tools. * Must be a U.S. citizen or lawful permanent resident (green card holder). * Must reside in the Washington, DC metropolitan area and be able to work on-site at Government offices. * Must be able to obtain and maintain a Public Trust background investigation., * Master's degree in a related field and 7+ years of data engineering experience, including support of federal agency programs. * Experience evaluating retrieval quality and building RAG pipelines with vector stores (e.g., pgvector, OpenSearch). * Experience with MLOps tooling, including model and dataset versioning, feature stores, or model registries. * Generative AI, LLM security, or data certification (e.g., Databricks Generative AI Engineer, Snowflake SnowPro, AWS, Microsoft Azure, or Google Cloud). * An active Public Trust or prior federal background investigation. ## Description * Design and maintain data ingestion, transformation, and processing pipelines (ETL/ELT) for AI training, evaluation, retrieval, and operations, including support for data migration and cleansing. * Curate, validate, and version datasets, and maintain dataset inventories, metadata, lineage, provenance, and ingestion logs. * Implement automated data-quality checks for duplication, schema changes, completeness, and freshness, and maintain dataset quality scorecards and drift reports. * Design data models, vector stores, and embedding schemas for Retrieval-Augmented Generation (RAG) knowledge bases, and re-index content when sources change. * Measure retrieval and model quality against established baselines using metrics such as precision/recall, MRR, NDCG, and context relevance. * Prepare data-related deliverables, including AI model cards, ML and AI pipeline documentation, RAG/AI Pipeline Evaluation Reports, data dictionaries and embedding schema documentation, and responses to Government data calls. * Build secure structured-data access for AI applications (e.g., natural-language-to-SQL with query validation and role-based authorization), and support dashboards and operational analytics. * Monitor data and ML pipelines, troubleshoot failures, and support root-cause analysis, while keeping all data in FedRAMP-authorized cloud regions with FIPS-validated encryption. * Support Responsible AI practices by preparing representative evaluation datasets, testing AI outputs for bias, accuracy, and hallucination, and documenting results to meet federal AI governance requirements. * Secure the AI data path, from source datasets and embeddings to prompts and logs, and support AI risk testing such as data poisoning. ## Related Videos - [Navigating the AI Revolution in Software Development](https://www.wearedevelopers.com/videos/1266-navigating-the-ai-revolution-in-software-development) - [Fully Orchestrating Databricks from Airflow](https://www.wearedevelopers.com/videos/336-fully-orchestrating-databricks-from-airflow) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Getting to Know Your Legacy (System) with AI-Driven Software Archeology](https://www.wearedevelopers.com/videos/1437-getting-to-know-your-legacy-system-with-ai-driven-software-archeology) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)