> Markdown version of [/jobs/ext/2203599-mid-level-data-engineer-130-006](https://www.wearedevelopers.com/jobs/ext/2203599-mid-level-data-engineer-130-006). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Mid-Level Data Engineer 130-006 - **Company:** Ic-cap Llc - **Location:** Alexandria, VA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Geographic Information Systems, Agile Methodology, Artificial Intelligence, Apache HTTP Server, ArcGIS (Software), Data Architecture, Information Engineering, Data Infrastructure, Extract Transform Load (ETL), Data Transformation, Data Retrieval, R (Programming Language), Graph Database, Integrated Development Environments, Data Intelligence, Python (Programming Language), NoSQL, Web Ontology Language, Software Engineering, SPARQL, SQL Databases, Tableau (Software), Enterprise Data Management, Scripting, Delivery Pipeline, Change Data Capture, Triple Store, Debezium, Kubernetes, Information Technology, Data Lineage, Apache Kafka, Apache Nifi, Graphql, Data Pipelines, Docker, Databricks - **Published:** August 23, 2026 - **Apply:** https://www.clearancejobs.com/jobs/9106111/mid-level-data-engineer-130-006 ## About the Role * Bachelor's degree in Computer Science, Software Engineering, or related STEM field AND 3-6 years of data engineering experience in classified IC or DoD environments * Proficiency in Python for ETL scripting, data transformation, and pipeline automation; familiarity with DIA baseline (Python, R, SQL, ArcGIS, Tableau) * Experience with Apache NiFi, Databricks, or equivalent enterprise data orchestration platforms * Proficiency in SQL; experience with NoSQL/graph databases; SPARQL query development for RDF data * Experience with SHACL constraint validation, data quality monitoring, and provenance tracking in classified data environments * Understanding of containerization (Docker, Kubernetes) and infrastructure-as-code principles per PWS §4.1.4 * Agile/SAFe software development environment experience Desired: * Direct experience with RDF triple stores: GraphDB Enterprise, Apache Jena Fuseki, or equivalent at production enterprise scale * Experience with OWL 2 reasoning, GeoSPARQL, or geospatial data engineering for spatially-enabled intelligence per PWS §4.1.4 Advanced Geospatial Analysis * Familiarity with DIEKM/DICO standards and DIA MARS or TALOS program data architecture * Experience with Debezium CDC, Kafka Streams, or real-time change data capture for schema change detection * Knowledge of PROV-O, SKOS, or other W3C provenance and vocabulary standards for data lineage * Experience with AgreementMakerLight, LogMap, or ontology matching tooling for source vocabulary alignment * Willingness and eligibility for COCOM deployment if activated per PWS §6.3 and §10 * SAFe Agile certification Security Clearance: * Active TS/SCI and the willingness to sit for a polygraph, if needed ## Description The Mid-Level Data Engineer designs, builds, and operates the scalable data pipelines, ingestion frameworks, metadata governance systems, and RDF triple store infrastructure that power DIA's enterprise MARS OBI platform. This role sits at the operational core, ensuring that high-quality, semantically consistent, provenance-tracked intelligence data flows reliably from authoritative sources into the analytic environment. All pipelines must integrate into DIA's IT environment. Duties may include: * Design, implement, and optimize scalable ETL/ELT data pipelines using Apache NiFi, Databricks, Apache Kafka, and Python to ingest, transform, normalize, and load multi-INT intelligence datasets into the MARS RDF triple store. * Develop and optimize data queries via multiple protocols explicitly: GraphQL, SPARQL, SHACL, and SQL - enabling semantic data retrieval and reasoning across knowledge graphs. * Develop and maintain SKOS-based semantic mapping registries per Ontology & Knowledge Modeling, aligning every source field and code value to approved OBI/DICO term URIs across all authoritative data sources. * Build and enforce SHACL validation shapes for cardinality, datatype, value range, and relationship constraints; execute fail-fast validation at ingestion and nightly bulk reconciliation across the full graph. * Manage PROV-O provenance tagging on all generated triples, maintaining complete data lineage from source record through transformation and graph load for every intelligence object - supporting AI documentation requirements. * Implement version-controlled GitOps deployment of SHACL shapes, mapping rules, ontology versions, and SPARQL CONSTRUCT queries per infrastructure-as-code and containerization requirements. * Operate and maintain enterprise RDF triple store infrastructure (GraphDB Enterprise, Apache Jena Fuseki) per secure storage and highly available distributed access requirements; monitor P95/P99 query latency. * Monitor pipeline health, ingestion throughput, SHACL validation pass rates, and provenance completeness metrics per Advanced Analytics & Modeling requirement; report status to quality dashboards. * Upon Government activation, provide Data Engineer support to assigned Combatant Command locations; maintain pipeline operations from COCOM environments. * Participate in SAFe ceremonies; contribute data engineering user stories and continuously deployed pipeline improvements at each Program Increment. ## Related Videos - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [NoSQL Data Modeling for Front-end Developers](https://www.wearedevelopers.com/videos/297-nosql-data-modeling-for-front-end-developers) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Dev Digest 134 - Where pixels sing?](https://www.wearedevelopers.com/magazine/477-dev-digest-134-where-pixels-sing) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know)