Data Engineer - Analytic Platform & Data Pipelines

Coleman Contracting, Inc.
Herndon, VA, United States
6 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$135,000.0 - $185,000.0
Working hours
Regular working hours

Tech stack

Query Performance Java (Programming Language) Geographic Information Systems Application Programming Interfaces (APIs) Airflow Amazon Web Services Data Analysis Microsoft Azure Cloud Computing Computer Programming Continuous Integration Data Validation
+38 more
Information Engineering Data Governance Extract Transform Load (ETL) Data Security Serialization Data Structures Data Stores Distributed Computing Environment Elasticsearch Graph Database JSON Python (Programming Language) PostgreSQL Metadata MongoDB NoSQL Open Source Technology Operational Databases PostGIS Role-Based Access Control Cloud Services Secure Coding SQL Databases Data Streaming Unstructured Data Workflow Management Systems Parquet Apache Spark Indexer Data Lakes Kubernetes Data Lineage Geospatial Data Abstraction Library (GDAL) Avro Dask Apache Kafka Data Pipelines Docker

Job description

Role Overview3GIMBALS is seeking a Data Engineer to design, build, and maintain the data pipelines and infrastructure that power our unclassified PAI/CAI-based analytic platform. This role is responsible for ingesting, transforming, and curating large volumes of structured and unstructured data from diverse open and commercial sources; building resilient, automated ETL/ELT workflows; and ensuring data is high-quality, well-governed, and analysis-ready for the downstream analytics, knowledge graph, and modeling teams. The ideal candidate is comfortable working with messy, multi-source data at scale within secure development environments.Key ResponsibilitiesData Pipeline Development & IngestionDesign and build scalable batch and streaming pipelines to ingest structured and unstructured data from PAI/CAI sources, APIs, and third-party feedsDevelop ETL/ELT workflows to normalize, enrich, and transform heterogeneous data into standardized schemasBuild and maintain automated ingestion

Requirements

connectors for web, document, geospatial, and tabular data sourcesManage data orchestration and scheduling using tools such as Airflow, Dagster, or PrefectData Modeling & StorageDesign and maintain data models, schemas, and storage layers across relational, NoSQL, and object storesBuild and maintain data lakes/lakehouses and curated, analysis-ready data martsOptimize partitioning, indexing, and query performance for large datasetsSupport entity resolution and data linking in coordination with the knowledge graph and modeling teamsData Quality, Governance & LineageImplement data validation, quality checks, and monitoring across pipelinesEstablish data lineage, cataloging, and metadata managementEnforce data governance, provenance tracking, and source attribution appropriate for PAI/CAI dataDocument datasets, schemas, and pipeline logic for downstream consumersSecurity & ComplianceEnsure pipelines and data stores meet security requirements for operation in sensitive environmentsImplement encryption, access control, and secure data-handling practicesSupport Authority to Operate (ATO) processes and compliance frameworksRequired QualificationsTechnical Expertise4+ years of data engineering experience building and operating production data pipelinesStrong programming skills in Python and SQL (Scala or Java a plus)Experience with distributed data processing frameworks (Spark, Dask, or similar)Hands-on experience with workflow orchestration tools (Airflow, Dagster, Prefect)Proficiency with relational and NoSQL databases (PostgreSQL, MongoDB, Elasticsearch, etc.)Experience with cloud data platforms and services (AWS, Azure, or GCP)Data & InfrastructureExperience designing data models, warehouses, and lakehouse architecturesFamiliarity with data formats and serialization (Parquet, Avro, JSON, GeoJSON)Understanding of data quality, lineage, and governance practicesExperience with containerization (Docker) and CI/CD for data workflowsDomain KnowledgeExperience working with large-scale, heterogeneous, or open-source datasetsUnderstanding of data provenance and source-attribution requirementsPreferred QualificationsActive security clearance or ability to obtain oneExperience in government, defense, or intelligence contracting environmentsFamiliarity with PAI/CAI (publicly and commercially available information) data sourcesExperience with geospatial data processing (PostGIS, GDAL, or similar)Knowledge of graph data structures and preparing data for knowledge graphsExperience with streaming platforms (Kafka, Kinesis)Familiarity with federal compliance frameworks (FedRAMP, FISMA, NIST 800-53)Technical EnvironmentLanguages: Python, SQL (Scala/Java a plus)Processing: Spark, Airflow/Dagster/Prefect, streaming frameworksStorage: PostgreSQL, Elasticsearch, object storage / data lake, ParquetInfrastructure: Docker, Kubernetes, cloud platforms (AWS GovCloud, Azure Government)Security: Encryption at rest and in transit, RBAC, secure data handling This role is central to the platform: the data engineering team delivers the clean, trustworthy, well-documented data that every analytic, knowledge graph, and risk-modeling capability depends on.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:02 min

Audience Q&A on data formats and engine tradeoffs

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

2:19 min

Scaling performance across multiple GPUs using specialized frameworks

Paul Graham Paul Graham · World Congress 2025

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

1:52 min

Customizing block storage tiers and formats

Ricardo Sueiras Sueiras · LIVE

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters · World Congress 2025

Videos

See all

Related articles

See all