> Markdown version of [/jobs/ext/2918986-data-engineer-analytic-platform-data-pipelines](https://www.wearedevelopers.com/jobs/ext/2918986-data-engineer-analytic-platform-data-pipelines). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - Analytic Platform & Data Pipelines - **Company:** Coleman Contracting, Inc. - **Location:** Herndon, VA, United States - **Salary:** $135,000.0 - $185,000.0 - **Contract:** Permanent contract - **Skills:** Query Performance, Java (Programming Language), Geographic Information Systems, Application Programming Interfaces (APIs), Airflow, Amazon Web Services, Data Analysis, Microsoft Azure, Cloud Computing, Computer Programming, Continuous Integration, Data Validation, Information Engineering, Data Governance, Extract Transform Load (ETL), Data Security, Serialization, Data Structures, Data Stores, Distributed Computing Environment, Elasticsearch, Graph Database, JSON, Python (Programming Language), PostgreSQL, Metadata, MongoDB, NoSQL, Open Source Technology, Operational Databases, PostGIS, Role-Based Access Control, Cloud Services, Secure Coding, SQL Databases, Data Streaming, Unstructured Data, Workflow Management Systems, Parquet, Apache Spark, Indexer, Data Lakes, Kubernetes, Data Lineage, Geospatial Data Abstraction Library (GDAL), Avro, Dask, Apache Kafka, Data Pipelines, Docker - **Published:** September 15, 2026 - **Apply:** https://www.careerjet.com/jobad/us538be712b221af03a48028b281e63411 ## About the Role connectors for web, document, geospatial, and tabular data sourcesManage data orchestration and scheduling using tools such as Airflow, Dagster, or PrefectData Modeling & StorageDesign and maintain data models, schemas, and storage layers across relational, NoSQL, and object storesBuild and maintain data lakes/lakehouses and curated, analysis-ready data martsOptimize partitioning, indexing, and query performance for large datasetsSupport entity resolution and data linking in coordination with the knowledge graph and modeling teamsData Quality, Governance & LineageImplement data validation, quality checks, and monitoring across pipelinesEstablish data lineage, cataloging, and metadata managementEnforce data governance, provenance tracking, and source attribution appropriate for PAI/CAI dataDocument datasets, schemas, and pipeline logic for downstream consumersSecurity & ComplianceEnsure pipelines and data stores meet security requirements for operation in sensitive environmentsImplement encryption, access control, and secure data-handling practicesSupport Authority to Operate (ATO) processes and compliance frameworksRequired QualificationsTechnical Expertise4+ years of data engineering experience building and operating production data pipelinesStrong programming skills in Python and SQL (Scala or Java a plus)Experience with distributed data processing frameworks (Spark, Dask, or similar)Hands-on experience with workflow orchestration tools (Airflow, Dagster, Prefect)Proficiency with relational and NoSQL databases (PostgreSQL, MongoDB, Elasticsearch, etc.)Experience with cloud data platforms and services (AWS, Azure, or GCP)Data & InfrastructureExperience designing data models, warehouses, and lakehouse architecturesFamiliarity with data formats and serialization (Parquet, Avro, JSON, GeoJSON)Understanding of data quality, lineage, and governance practicesExperience with containerization (Docker) and CI/CD for data workflowsDomain KnowledgeExperience working with large-scale, heterogeneous, or open-source datasetsUnderstanding of data provenance and source-attribution requirementsPreferred QualificationsActive security clearance or ability to obtain oneExperience in government, defense, or intelligence contracting environmentsFamiliarity with PAI/CAI (publicly and commercially available information) data sourcesExperience with geospatial data processing (PostGIS, GDAL, or similar)Knowledge of graph data structures and preparing data for knowledge graphsExperience with streaming platforms (Kafka, Kinesis)Familiarity with federal compliance frameworks (FedRAMP, FISMA, NIST 800-53)Technical EnvironmentLanguages: Python, SQL (Scala/Java a plus)Processing: Spark, Airflow/Dagster/Prefect, streaming frameworksStorage: PostgreSQL, Elasticsearch, object storage / data lake, ParquetInfrastructure: Docker, Kubernetes, cloud platforms (AWS GovCloud, Azure Government)Security: Encryption at rest and in transit, RBAC, secure data handling This role is central to the platform: the data engineering team delivers the clean, trustworthy, well-documented data that every analytic, knowledge graph, and risk-modeling capability depends on. ## Description Role Overview3GIMBALS is seeking a Data Engineer to design, build, and maintain the data pipelines and infrastructure that power our unclassified PAI/CAI-based analytic platform. This role is responsible for ingesting, transforming, and curating large volumes of structured and unstructured data from diverse open and commercial sources; building resilient, automated ETL/ELT workflows; and ensuring data is high-quality, well-governed, and analysis-ready for the downstream analytics, knowledge graph, and modeling teams. The ideal candidate is comfortable working with messy, multi-source data at scale within secure development environments.Key ResponsibilitiesData Pipeline Development & IngestionDesign and build scalable batch and streaming pipelines to ingest structured and unstructured data from PAI/CAI sources, APIs, and third-party feedsDevelop ETL/ELT workflows to normalize, enrich, and transform heterogeneous data into standardized schemasBuild and maintain automated ingestion ## Related Videos - [From event streaming to event sourcing 101](https://www.wearedevelopers.com/videos/91-from-event-streaming-to-event-sourcing-101) - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Introducing JSON Structure](https://www.wearedevelopers.com/videos/100219-introducing-json-structure) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Green Cloud Computing](https://www.wearedevelopers.com/videos/592-green-cloud-computing) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)