Data Engineer (AWS, Spark)

Flex Ltd.
Washington, DC, United States
1 day ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Compensation
$103,000.0 - $140,000.0
Working hours
Regular working hours

Tech stack

Microsoft Excel Artificial Intelligence Amazon Web Services Amazon S3 Apache HTTP Server Information Engineering Extract Transform Load (ETL) IBM InfoSphere DataStage Programming Tools Amazon DynamoDB Python (Programming Language) Key Management
+16 more
PostgreSQL Machine Learning Microsoft PowerPoint SQL Databases Parquet Delivery Pipeline Apache Spark Cloudformation Pyspark Storage Technologies Information Technology Data Lineage Amazon Simple Queue Service (SQS) Terraform Data Pipelines Amazon Elastic Mapreduce (EMR)

Job description

Build the pipelines that move a federal agency’s data from source to platform. In this role, you will work with a premier data and technology innovation hub and certified Benefit Corporation operating at the center of the US federal government’s mission. They don’t hire just to fill seats-they hire top-tier technical talent and move their best people to where the hardest problems are. As the work evolves, you will learn new systems and tools, take on greater responsibility, and help develop internal capabilities and new lines of business. Your First Project Your first project will have you building ingest, processing, and storage architecture at scale. Depending on the specific assignment, you will: Build Spark-based extract, transform, and load (ETL) pipelines using Glue, Amazon EMR, Lambda, and Step Functions. Write scalable data processing workflows in Python and PySpark. Design S3 data-lake layers (including Parquet, partitioning, and lifecycle policies) feeding Apache Iceberg tables. Connect PostgreSQL on Amazon Aurora and DynamoDB, using Trino for federated Structured Query Language (SQL) across relational stores and the lake. Deliver automated pipelines that keep data clean, trustworthy, and audit-ready at high-volume federal scale. (Note: This is just where you start, not the final shape of your career here.) Who You Are You are a data engineer driven to continuously improve, and you know which areas you want to master next. You care as much about whether data is trustworthy as whether it arrives on time, and you do your best work alongside teammates who push you. You experiment, learn, and iterate quickly, preferring to own an outcome rather than just be handed a task. What You Bring to the Table (Requirements), Machine Learning Engineer 4 (Python, AWS, SQL, GenAI) (Enterprise Platforms Technology) Do you love building and pioneering in the AI and technology space? Do you enjoy solving com…

  • 1 day ago

Requirements

Clearance & Location: Sole United States citizenship and the ability to obtain a Public Trust determination are required. This is a full-time W-2 hybrid role based in the Washington, DC metropolitan area. Experience & Education: 4+ years of dedicated data engineering experience and a Bachelor’s degree. Technical Stack: Hands-on Spark ETL on AWS (Glue, Amazon EMR) using Python and PySpark. S3 data-lake design (Parquet, partitioning, lifecycle) feeding Apache Iceberg tables, Amazon Aurora PostgreSQL, and DynamoDB. Event orchestration (Lambda, Step Functions, SQS/SNS) with secrets management and monitoring. Data quality, validation, data lineage, and Infrastructure-as-Code (CloudFormation or Terraform). Basic proficiency in writing, PowerPoint, and Excel. Bonus Points (Preferred) Master’s degree in a relevant field. Experience with Trino or comparable federated SQL across lake and relational stores. Experience with Apache Ranger-governed access controls. Legacy ETL migration experience (e.g., DataStage). Federal IT experience or high-volume data experience, alongside familiarity with AI-assisted developer tooling., Machine Learning Engineer 5 (IC) Do you love building and pioneering in the AI and technology space? Do you enjoy solving complex business problems in a fast-paced, collaborative…

Benefits & conditions

Base Salary: $103,000 to $140,000 per year. Health Coverage: 100% employer-paid medical, dental, and vision insurance for employees (50% coverage for dependents). Retirement & Security: 401(k) matched 100% up to 4% (vesting immediately); employer-paid life, accidental death, and short/long-term disability insurance. Growth & Time Off: Unlimited paid time off, extensive onboarding, sponsored technical certifications (such as AWS and DCAM), and continuing education.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann Chris Heilmann +3 · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:50 min

How Parquet metadata enables efficient data reading

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:44 min

Automating storage savings with S3 intelligent tiering

Sébastien Stormacq · World Congress 2021

Videos

See all

Related articles

See all