> Markdown version of [/jobs/ext/3140798-aws-lakehouse-data-engineer](https://www.wearedevelopers.com/jobs/ext/3140798-aws-lakehouse-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AWS Lakehouse Data Engineer - **Company:** System One - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Microsoft Access, Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Amazon S3, Apache HTTP Server, Automation of Tests, Information Engineering, Data Governance, Extract Transform Load (ETL), Data Visualization, Relational Databases, Cursor (Graphical User Interface Elements), Distributed Data Store, Github, Identity and Access Management, Python (Programming Language), Key Management, Network Security, Machine Learning, Metadata Repositories, Performance Tuning, Systems Development Life Cycle, Query Optimization, Role-Based Access Control, Power BI, SQL Databases, Data Streaming, Tableau (Software), Parquet, Data Logging, Data Classification, Sql Optimization, GitHub Copilot, Delivery Pipeline, Change Data Capture, Infrastructure as Code (IaC), Git, Cloudformation, Data Lakes, Pyspark, Information Technology, Data Lineage, AWS Glue, Data Management, Terraform, GPT, Data Pipelines, Amazon Elastic Mapreduce (EMR), Docker, Jenkins, Amazon Redshift, Databricks - **Published:** September 29, 2026 - **Apply:** https://www.dice.com/job-detail/ebe0b23b-f6bc-4541-9f02-b066d994636d ## About the Role * Bachelor's degree in Engineering, Information Technology, Computer Science, Data Engineering, or related field, or four (4) years of equivalent practical experience. * Six (6) years of relevant experience. * Hands-on experience building AWS-native data lake or lakehouse architectures on Amazon S3. * Strong experience developing production ETL/ELT pipelines with Python and PySpark, including data modeling, transformation, and performance tuning. * Hands-on experience with Apache Iceberg, including ACID transactions, schema evolution, time travel, and query optimization. * Advanced SQL skills supporting analytical workloads, reporting, and data visualization. * Proven experience with data governance, cataloging, lineage, and access control using AWS services. * Knowledge of AWS security fundamentals: IAM, KMS, secrets management, network security, logging, SDLC. * Proven experience with Infrastructure as Code (IaC) and operating data platforms across environments. * Experience with CI/CD pipelines for data workflows with testing, deployment, environment promotion, and rollback. * Troubleshooting distributed data workloads, performance optimization, and cost management skills. * Excellent collaboration and communication skills to coordinate with cross-team stakeholders. Would Be Nice to Have * Experience with Databricks, Delta Lake, migrating workloads to AWS-native services, and Apache Iceberg. * Familiarity with AWS Step Functions, MWAA, Kinesis, DMS, Lambda, MSK, or similar services. * Experience with modern DevOps tools: Git, Terraform, CloudFormation, Jenkins, CodePipeline, GitHub Actions, Docker. * Knowledge of BI and visualization tools like Amazon QuickSight, Tableau, Power BI. * Familiarity with AI-assisted coding tools such as GitHub Copilot, ChatGPT, Cursor, or Kiro. * Knowledge of graph modeling, ontology, taxonomy, entity resolution, and hybrid retrieval techniques. ## Description * Build and operate data pipelines (batch and streaming) from APIs, relational databases, file drops, event streams, and external partners. * Design, implement, test, and optimize ETL/ELT pipelines using Python and PySpark to produce analytics-ready datasets for reporting, visualization, and machine learning. * Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks. * Improve pipeline reliability through automated testing, orchestration, monitoring, retries, and operational runbooks. * Design and implement a Delta Lakehouse-style data platform on AWS using native services to provide Databricks-like capabilities for data engineering, analysis, and data visualization. * Build and manage a scalable lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet. * Implement SQL-like table reliability features including ACID transactions, schema evolution, snapshot isolation, and time travel using Apache Iceberg. * Enable fast, interactive queries of lakehouse data via AWS-native services like Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift where appropriate. * Optimize performance and cost through partitioning, file sizing, caching, lifecycle policies, and separating compute from storage. * Establish standardized environments for development, testing, and production with consistent configuration and controlled promotion. * Implement data governance, access control, lineage, and quality measures utilizing AWS-native services including AWS Lake Formation, AWS Glue Data Catalog, IAM, KMS. * Create a metadata repository with cataloging, ownership, classification, tagging, and discoverability features. * Enable end-to-end data lineage for audit and regulatory compliance. * Apply policy-based access, least privilege, data classification, retention, encryption, and secure handling controls. * Build data quality checks for freshness, completeness, validity, and anomaly detection, and publish SLA/SLO metrics. * Automate AWS provisioning with Infrastructure as Code (IaC), develop CI/CD pipelines for data components, and ensure platform observability. * Work collaboratively with cross-functional teams and maintain high-quality engineering documentation., System One, and its subsidiaries including Joulé and Mountain Ltd., are leaders in delivering outsourced services and workforce solutions across North America. We help clients get work done more efficiently and economically, without compromising quality. System One not only serves as a valued partner for our clients, but we offer eligible employees health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, voluntary plans, as well as participation in a 401(k) plan. ## Related Videos - [Livecoding with AI](https://www.wearedevelopers.com/videos/1201-livecoding-with-ai) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [ Evaluating AI models for code comprehension](https://www.wearedevelopers.com/videos/1462-evaluating-ai-models-for-code-comprehension) - [OLTP in the Lakehouse: Redefining Data for AI Workloads](https://www.wearedevelopers.com/videos/2038-oltp-in-the-lakehouse-redefining-data-for-ai-workloads) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)