Data Engineer, AWS Data Lake

D9tech Resources LLC
United States
16 days ago
Apply on www.clearancejobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Part-time (≤ 32 hours)
Working hours
Regular working hours

Tech stack

Amazon Web Services Amazon S3 CompTIA Security+ Continuous Integration Extract Transform Load (ETL) Data Transformation Python (Programming Language) Raw Data Data Processing SC Clearance Data Lakes Pyspark
+5 more
Storage Technologies AWS Glue Data Analytics AWS Data Analytics Terraform

Job description

D9Tech Resources is sourcing a cleared Data Engineer to build and sustain a data lake on AWS. The work sits where storage, movement, and governance meet: you will land raw data, shape it into something analysts can actually query, and keep the pipelines that do it running clean. Ingestion is only half the job. The other half is making the environment repeatable, so that infrastructure ships as code and compliance evidence is produced by the pipeline rather than reconstructed after the fact. This is a hands-on engineering seat, not an advisory one. The right candidate has written Glue jobs that failed, debugged them, and shipped the fix. If you are comfortable in a Terraform module and equally comfortable explaining a schema decision to a data consumer who does not care how it was built, this posting was written for you.

WHAT YOU WILL DO

  • Build and operate the data lake. Design S3-backed storage layers, partitioning strategies, and catalog structure that keep query cost and query time under control as volume grows.
  • Own the ETL. Develop, schedule, and tune AWS Glue jobs and crawlers, and support the surrounding data processing workflow end to end, from source ingestion through curated output.
  • Ship infrastructure as code. Write and maintain Terraform for the lake, the pipelines, and the supporting AWS services, keeping environments consistent and changes reviewable.
  • Automate compliance. Codify security and configuration requirements so controls are enforced and evidenced automatically instead of checked by hand.
  • Guard data quality. Build validation, monitoring, and alerting into the pipelines, then act on what they surface.
  • Document what you build. Maintain data flow diagrams, runbooks, and schema documentation that survive your absence.
  • Partner across the mission. Work with analysts, application teams, and security stakeholders to translate data requirements into working pipelines.

Requirements

  • Active Secret clearance.
  • U.S. citizenship.
  • Hands-on experience designing, building, or operating a data lake in AWS, including S3 storage design, partitioning, and cataloging.
  • Demonstrated ETL and data processing depth with AWS Glue, including job development, crawlers, and the AWS Glue Data Catalog.
  • Working proficiency with Terraform for provisioning and maintaining AWS infrastructure.
  • Python or PySpark for data transformation work.
  • Comfort operating independently in a fully remote, distributed team., * Compliance as Code experience: policy, control, and configuration enforcement expressed and validated through automation.
  • Additional AWS analytics services such as Athena, Lake Formation, Redshift, EMR, Step Functions, Kinesis, or Lambda.
  • CI/CD pipeline experience for data and infrastructure deployments.
  • AWS certification such as Data Engineer Associate, Solutions Architect Associate, or Data Analytics Specialty.
  • Prior delivery on a Federal or Department of Defense program, and familiarity with NIST SP 800-53 or DoD security requirements.
  • CompTIA Security+ (Sec+ CE) or an equivalent DoD 8140 baseline certification.

Benefits & conditions

The role is fully remote within the continental United States. Expected effort is approximately 32 hours per week. Because the work supports a cleared program, the selected engineer must hold and maintain an active Secret clearance for the duration of the assignment, and clearance eligibility is verified before any offer is extended.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.clearancejobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

4:09 min

Challenges of interpreting raw data with language models

Clemens Vasters Clemens Vasters · World Congress 2025

55 sec

Validating data processing architectures via containerized events

Modood Alvi · World Congress 2025

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:09 min

Uncovering hidden coordinate manipulation communities in binary data

Nolan Royalty · Coffee With Developers

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all