AWS Data Engineer

The Joule
McLean, VA, United States
about 1 month ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Compensation
$150,000.0 - $170,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Amazon S3 Data Analysis Apache HTTP Server Automation of Tests Information Engineering Data Governance Extract Transform Load (ETL) Data Security Data Visualization
+27 more
Relational Databases Distributed Data Store Identity and Access Management Python (Programming Language) Key Management Network Security Machine Learning Metadata Repositories Operational Data Store Performance Tuning Query Optimization Runbook SQL Databases Parquet Cloud Platform System Data Classification Sql Optimization Change Data Capture Infrastructure as Code (IaC) Data Lakes Pyspark Information Technology Data Lineage AWS Glue Data Pipelines Amazon Elastic Mapreduce (EMR) Amazon Redshift

Job description

  • Build and operate data pipelines (batch and streaming) from various sources including APIs, relational databases, file drops, event streams, and external partners.
  • Design, implement, and optimize ETL/ELT pipelines using Python and PySpark to produce analytics-ready datasets for reporting, visualization, and machine learning.
  • Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks.
  • Improve pipeline reliability through automated testing, orchestration, monitoring, retry handling, and operational runbooks.
  • Build and manage a scalable lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet.
  • Implement SQL-like table reliability features including ACID transactions, schema evolution, partition evolution, snapshot isolation, time travel, and rollback capabilities.
  • Enable fast, interactive querying of lakehouse data using AWS-native query and compute services such as Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift.
  • Optimize performance and cost efficiency through partitioning, compaction, file sizing, caching, lifecycle policies, and efficient compute/storage separation.
  • Establish standardized development, testing, and production environments with consistent configuration and controlled promotion across stages.
  • Implement data governance and fine-grained access control utilizing AWS-native services like AWS Lake Formation, AWS Glue Data Catalog, IAM, KMS, and related security tools.
  • Create a managed metadata repository for dataset cataloging, ownership, tagging, classification, and discoverability.
  • Support end-to-end data lineage for source, transformation, and consumption to facilitate auditability and impact analysis.
  • Apply security policies such as least privilege access, data classification, encryption, retention, and secure data handling.
  • Build operational data quality checks for metrics such as freshness, completeness, validity, and anomaly detection, along with publishing SLAs/SLOs.
  • Implement automated AWS provisioning through Infrastructure as Code (IaC) to ensure consistent, secure environments.
  • Enhance CI/CD pipelines for data workflows and lakehouse components, including automated testing, security validation, packaging, deployment, promotion, and rollback.
  • Maintain observability with centralized metrics, logs, traces, alerts, dashboards, runbooks, and incident response procedures.
  • Continually evaluate platform performance, scalability, reliability, security, and cost, and implement measurable improvements.
  • Collaborate closely with data, application, analytics, AI/ML, security, networking, and cloud platform teams to support mission-critical requirements.
  • Maintain high-quality documentation including architecture diagrams, SOPs, data models, interface specs, and operational runbooks.
  • Present technical findings, trade-offs, risks, and recommendations clearly to stakeholders.

Requirements

  • Bachelor’s degree in Engineering, Information Technology, Computer Science, Data Engineering, or a related field, or four (4) years of equivalent practical experience.
  • Six (6) years of relevant hands-on experience in data engineering.
  • Extensive experience designing, implementing, and operating AWS-native data lake or lakehouse architectures with Amazon S3, AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift.
  • Proven ability developing production ETL/ELT pipelines using Python and PySpark, including data modeling, transformation, tuning, and error handling.
  • Hands-on experience with Apache Iceberg covering ACID transactions, snapshots, schema evolution, and query optimization.
  • Advanced SQL skills supporting analytical queries, reporting, and data visualization workloads.
  • Demonstrated experience with data governance, cataloging, lineage, ownership, classification, and access controls.
  • Knowledge of AWS security fundamentals including IAM, encryption (KMS), secrets management, and network security.
  • Experience provisioning resources via Infrastructure as Code (IaC) and managing multi-environment platforms.
  • Skilled in building and maintaining CI/CD pipelines for data workflows with automation, testing, and rollback strategies.
  • Strong troubleshooting skills for distributed data workloads with focus on performance, reliability, and cost management.

About the company

System One, and its subsidiaries including Joulé and Mountain Ltd., are leaders in delivering outsourced services and workforce solutions across North America. We help clients get work done more efficiently and economically, without compromising quality. System One not only serves as a valued partner for our clients, but we offer eligible employees health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, voluntary plans, as well as participation in a 401(k) plan.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:50 min

How Parquet metadata enables efficient data reading

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

2:50 min

Introduction and the value of runbooks

Hila Fish · World Congress 2023

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

1:43 min

AWS infrastructure stack and data flow pipeline overview

Artem Volk Artem Volk +1 · World Congress 2024

2:03 min

Introduction to open table formats built on Parquet

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

Videos

See all

Related articles

See all