Data Engineer (AWS, Spark)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+16 more
Job description
Build the pipelines that move a federal agency’s data from source to platform. In this role, you will work with a premier data and technology innovation hub and certified Benefit Corporation operating at the center of the US federal government’s mission. They don’t hire just to fill seats-they hire top-tier technical talent and move their best people to where the hardest problems are. As the work evolves, you will learn new systems and tools, take on greater responsibility, and help develop internal capabilities and new lines of business. Your First Project Your first project will have you building ingest, processing, and storage architecture at scale. Depending on the specific assignment, you will: Build Spark-based extract, transform, and load (ETL) pipelines using Glue, Amazon EMR, Lambda, and Step Functions. Write scalable data processing workflows in Python and PySpark. Design S3 data-lake layers (including Parquet, partitioning, and lifecycle policies) feeding Apache Iceberg tables. Connect PostgreSQL on Amazon Aurora and DynamoDB, using Trino for federated Structured Query Language (SQL) across relational stores and the lake. Deliver automated pipelines that keep data clean, trustworthy, and audit-ready at high-volume federal scale. (Note: This is just where you start, not the final shape of your career here.) Who You Are You are a data engineer driven to continuously improve, and you know which areas you want to master next. You care as much about whether data is trustworthy as whether it arrives on time, and you do your best work alongside teammates who push you. You experiment, learn, and iterate quickly, preferring to own an outcome rather than just be handed a task. What You Bring to the Table (Requirements), Machine Learning Engineer 4 (Python, AWS, SQL, GenAI) (Enterprise Platforms Technology) Do you love building and pioneering in the AI and technology space? Do you enjoy solving com…
- 1 day ago
Requirements
Clearance & Location: Sole United States citizenship and the ability to obtain a Public Trust determination are required. This is a full-time W-2 hybrid role based in the Washington, DC metropolitan area. Experience & Education: 4+ years of dedicated data engineering experience and a Bachelor’s degree. Technical Stack: Hands-on Spark ETL on AWS (Glue, Amazon EMR) using Python and PySpark. S3 data-lake design (Parquet, partitioning, lifecycle) feeding Apache Iceberg tables, Amazon Aurora PostgreSQL, and DynamoDB. Event orchestration (Lambda, Step Functions, SQS/SNS) with secrets management and monitoring. Data quality, validation, data lineage, and Infrastructure-as-Code (CloudFormation or Terraform). Basic proficiency in writing, PowerPoint, and Excel. Bonus Points (Preferred) Master’s degree in a relevant field. Experience with Trino or comparable federated SQL across lake and relational stores. Experience with Apache Ranger-governed access controls. Legacy ETL migration experience (e.g., DataStage). Federal IT experience or high-volume data experience, alongside familiarity with AI-assisted developer tooling., Machine Learning Engineer 5 (IC) Do you love building and pioneering in the AI and technology space? Do you enjoy solving complex business problems in a fast-paced, collaborative…
Benefits & conditions
Base Salary: $103,000 to $140,000 per year. Health Coverage: 100% employer-paid medical, dental, and vision insurance for employees (50% coverage for dependents). Retirement & Security: 401(k) matched 100% up to 4% (vesting immediately); employer-paid life, accidental death, and short/long-term disability insurance. Growth & Time Off: Unlimited paid time off, extensive onboarding, sponsored technical certifications (such as AWS and DCAM), and continuing education.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Data Engineer Salary UK
Making Data Warehouses Fast: A Developer’s Story
How to Become an AI Engineer
Dev Digest 120 - Apple and peers