> Markdown version of [/jobs/ext/2941673-data-engineer](https://www.wearedevelopers.com/jobs/ext/2941673-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Groundswell, LLC - **Location:** Tysons, VA, United States - **Experience:** Expert - **Salary:** $89,886.0 - $175,444.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Airflow, Amazon Web Services, Data Analysis, Audit Trail, Big Data, Computer Engineering, Continuous Integration, Information Engineering, Data Integration, Data Integrity, Extract Transform Load (ETL), Data Security, Distributed Data Store, R (Programming Language), JSON, Python (Programming Language), Key Management, Machine Learning, Operational Data Store, Query Optimization, Runbook, Extensible Markup Language (XML), Jupyter Notebook, Parquet, Data Ingestion, Large Language Models, Apache Spark, Information Technology, Data Lineage, Deployment Automation, AWS Glue, Data Management, Machine Learning Operations, Api Design, Data Pipelines, Databricks - **Published:** September 16, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/88331596/1 ## About the Role * Bachelor's degree in Computer Science, Computer Engineering, Mathematics, Statistics, or a related technical field. * 5+ years of professional data engineering experience, including production pipeline development and operations. * Strong SQLexpertise, including query optimization, data modeling, joins, window functions, and analysis of large datasets. * Proficiencyin Python, Java,R, or otherprogramming language used for data engineering. * Experience with modern data-integration and workflow frameworks such as Apache Airflow, AWS Glue, Spark, or comparable tools. * Experience with Jupyter Notebook or otherequvialenttools for analyzing data. * Experience with Python libraries such asnumpy, panda, and other libraries like this. * Hands-on experience with at least one major cloud provider, with AWS experience strongly preferred. * Experience implementing data-quality frameworks, validation processes, monitoring, and alerting for production pipelines. * Working knowledge of data security, access controls, encryption, auditability, and governance in a regulated, restricted, or compliance-oriented environment. * Experience working with APIs and structured data formats such as JSON, XML, CSV, and Parquet. * Ability to document technical decisions and collaborate with multidisciplinary teams and customer stakeholders. * Must be a U.S. Citizen per contract requirements. * Must be able to obtain andmaintaina Public Trust Clearancein accordance withcontract requirements. Preferred Qualifications * Experience building preprocessing, chunking, filtering, metadata-enrichment, or evaluation pipelines for LLM and other AI/ML workflows. * Experience operating data platforms with segmented networks, limited connectivity, strict change control, or formal authorization requirements. * Extra consideration given to those with C3.ai experience. * Active Public Trust Clearance., * AWSData Engineer Certification, Databricks, cloud data engineering, or other relevant professional certification is preferred. ## Description Groundswell is seeking an experienced Data Engineer to build secure, reliable data capabilities for a mission-focused platform that connects authoritative sources, operational data, and analytical services. This role spans data ingestion, integration, quality, governance, and pipeline operations. You will help make complex data usable and trustworthy by designing repeatable workflows that preserve provenance, enforce access controls, and support decision-ready applications and AI-enabled capabilities in controlled environments. What You'll Do * Design, develop, andoperatebatch and streaming ETL/ELT pipelines that ingest data from multiple structured and semi-structured sources into secure cloud data platforms. * Onboard new data sources by defining schemas, mappings, interfaces, validation rules, ownership, and operational support procedures. * Build automated data-quality checks for completeness, accuracy, consistency, timeliness, and referential integrity, with actionable monitoring and alerting. * Create data transformations and processing workflows that support operational applications, analytics, machine learning, and retrieval or inference workflows. * Preserve data lineage, provenance, version history, audit trails, and handling metadata throughout the data lifecycle so users can understand where data came from and how it was changed. * Implement secure data practices, including least-privilege access, encryption, secrets management, privacy protections, and controlsappropriate forsensitive or classified information. * Develop andmaintaindata models, curated datasets, and service interfaces that provide consistent, well-documented access to authoritative information. * Automate deployments and pipeline operations using Infrastructure as Code, CI/CD, workflow orchestration, and environment-specific configuration. * Optimizedata pipelines for scalability, resiliency, performance, and cost across cloud platforms, including effective partitioning, parallelism, storage, and computeutilization. * Troubleshoot failures across distributed data systems using logs, metrics, lineage, and operational signals; communicate root causes and recovery plans clearly. * Collaborate with software engineers, cloud engineers, data scientists, security teams, architects, and customer stakeholders to translate mission needs into measurable data products. * Produce maintainable data contracts, architecture documentation, runbooks, test plans, and operational guidance for the broader team. * Contribute technicalexpertiseto solution planning and proposal efforts involving secure data platforms and AI-enabled mission capabilities. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Bridging AI and Nomad: a Go-based MCP Server for Cluster Control](https://www.wearedevelopers.com/videos/2063-bridging-ai-and-nomad-a-go-based-mcp-server-for-cluster-control) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story)