> Markdown version of [/jobs/ext/3455546-aws-data-engineer](https://www.wearedevelopers.com/jobs/ext/3455546-aws-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AWS Data Engineer - **Company:** The Joule - **Location:** McLean, VA, United States (Remote available) - **Experience:** Expert - **Salary:** $150,000.0 - $170,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Amazon S3, Data Analysis, Apache HTTP Server, Automation of Tests, Information Engineering, Data Governance, Extract Transform Load (ETL), Data Security, Data Visualization, Relational Databases, Distributed Data Store, Identity and Access Management, Python (Programming Language), Key Management, Network Security, Machine Learning, Metadata Repositories, Operational Data Store, Performance Tuning, Query Optimization, Runbook, SQL Databases, Parquet, Cloud Platform System, Data Classification, Sql Optimization, Change Data Capture, Infrastructure as Code (IaC), Data Lakes, Pyspark, Information Technology, Data Lineage, AWS Glue, Data Pipelines, Amazon Elastic Mapreduce (EMR), Amazon Redshift - **Published:** September 1, 2026 - **Apply:** https://www.dice.com/job-detail/a0dcc300-d36e-4e8f-8507-11c39ccaae35 ## About the Role * Bachelor's degree in Engineering, Information Technology, Computer Science, Data Engineering, or a related field, or four (4) years of equivalent practical experience. * Six (6) years of relevant hands-on experience in data engineering. * Extensive experience designing, implementing, and operating AWS-native data lake or lakehouse architectures with Amazon S3, AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift. * Proven ability developing production ETL/ELT pipelines using Python and PySpark, including data modeling, transformation, tuning, and error handling. * Hands-on experience with Apache Iceberg covering ACID transactions, snapshots, schema evolution, and query optimization. * Advanced SQL skills supporting analytical queries, reporting, and data visualization workloads. * Demonstrated experience with data governance, cataloging, lineage, ownership, classification, and access controls. * Knowledge of AWS security fundamentals including IAM, encryption (KMS), secrets management, and network security. * Experience provisioning resources via Infrastructure as Code (IaC) and managing multi-environment platforms. * Skilled in building and maintaining CI/CD pipelines for data workflows with automation, testing, and rollback strategies. * Strong troubleshooting skills for distributed data workloads with focus on performance, reliability, and cost management. ## Description * Build and operate data pipelines (batch and streaming) from various sources including APIs, relational databases, file drops, event streams, and external partners. * Design, implement, and optimize ETL/ELT pipelines using Python and PySpark to produce analytics-ready datasets for reporting, visualization, and machine learning. * Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks. * Improve pipeline reliability through automated testing, orchestration, monitoring, retry handling, and operational runbooks. * Build and manage a scalable lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet. * Implement SQL-like table reliability features including ACID transactions, schema evolution, partition evolution, snapshot isolation, time travel, and rollback capabilities. * Enable fast, interactive querying of lakehouse data using AWS-native query and compute services such as Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift. * Optimize performance and cost efficiency through partitioning, compaction, file sizing, caching, lifecycle policies, and efficient compute/storage separation. * Establish standardized development, testing, and production environments with consistent configuration and controlled promotion across stages. * Implement data governance and fine-grained access control utilizing AWS-native services like AWS Lake Formation, AWS Glue Data Catalog, IAM, KMS, and related security tools. * Create a managed metadata repository for dataset cataloging, ownership, tagging, classification, and discoverability. * Support end-to-end data lineage for source, transformation, and consumption to facilitate auditability and impact analysis. * Apply security policies such as least privilege access, data classification, encryption, retention, and secure data handling. * Build operational data quality checks for metrics such as freshness, completeness, validity, and anomaly detection, along with publishing SLAs/SLOs. * Implement automated AWS provisioning through Infrastructure as Code (IaC) to ensure consistent, secure environments. * Enhance CI/CD pipelines for data workflows and lakehouse components, including automated testing, security validation, packaging, deployment, promotion, and rollback. * Maintain observability with centralized metrics, logs, traces, alerts, dashboards, runbooks, and incident response procedures. * Continually evaluate platform performance, scalability, reliability, security, and cost, and implement measurable improvements. * Collaborate closely with data, application, analytics, AI/ML, security, networking, and cloud platform teams to support mission-critical requirements. * Maintain high-quality documentation including architecture diagrams, SOPs, data models, interface specs, and operational runbooks. * Present technical findings, trade-offs, risks, and recommendations clearly to stakeholders. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Parquet, Delta, Iceberg & Ducklake - An introduction for developers](https://www.wearedevelopers.com/videos/100075-parquet-delta-iceberg-ducklake-an-introduction-for-developers) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)