AWS Lakehouse Data Engineer

System One
United States
5 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours
Job source

Tech stack

Microsoft Access Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Amazon S3 Apache HTTP Server Automation of Tests Information Engineering Data Governance Extract Transform Load (ETL) Data Visualization Relational Databases
+41 more
Cursor (Graphical User Interface Elements) Distributed Data Store Github Identity and Access Management Python (Programming Language) Key Management Network Security Machine Learning Metadata Repositories Performance Tuning Systems Development Life Cycle Query Optimization Role-Based Access Control Power BI SQL Databases Data Streaming Tableau (Software) Parquet Data Logging Data Classification Sql Optimization GitHub Copilot Delivery Pipeline Change Data Capture Infrastructure as Code (IaC) Git Cloudformation Data Lakes Pyspark Information Technology Data Lineage AWS Glue Data Management Terraform GPT Data Pipelines Amazon Elastic Mapreduce (EMR) Docker Jenkins Amazon Redshift Databricks

Job description

  • Build and operate data pipelines (batch and streaming) from APIs, relational databases, file drops, event streams, and external partners.
  • Design, implement, test, and optimize ETL/ELT pipelines using Python and PySpark to produce analytics-ready datasets for reporting, visualization, and machine learning.
  • Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks.
  • Improve pipeline reliability through automated testing, orchestration, monitoring, retries, and operational runbooks.
  • Design and implement a Delta Lakehouse-style data platform on AWS using native services to provide Databricks-like capabilities for data engineering, analysis, and data visualization.
  • Build and manage a scalable lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet.
  • Implement SQL-like table reliability features including ACID transactions, schema evolution, snapshot isolation, and time travel using Apache Iceberg.
  • Enable fast, interactive queries of lakehouse data via AWS-native services like Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift where appropriate.
  • Optimize performance and cost through partitioning, file sizing, caching, lifecycle policies, and separating compute from storage.
  • Establish standardized environments for development, testing, and production with consistent configuration and controlled promotion.
  • Implement data governance, access control, lineage, and quality measures utilizing AWS-native services including AWS Lake Formation, AWS Glue Data Catalog, IAM, KMS.
  • Create a metadata repository with cataloging, ownership, classification, tagging, and discoverability features.
  • Enable end-to-end data lineage for audit and regulatory compliance.
  • Apply policy-based access, least privilege, data classification, retention, encryption, and secure handling controls.
  • Build data quality checks for freshness, completeness, validity, and anomaly detection, and publish SLA/SLO metrics.
  • Automate AWS provisioning with Infrastructure as Code (IaC), develop CI/CD pipelines for data components, and ensure platform observability.
  • Work collaboratively with cross-functional teams and maintain high-quality engineering documentation., System One, and its subsidiaries including Joulé and Mountain Ltd., are leaders in delivering outsourced services and workforce solutions across North America. We help clients get work done more efficiently and economically, without compromising quality. System One not only serves as a valued partner for our clients, but we offer eligible employees health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, voluntary plans, as well as participation in a 401(k) plan.

Requirements

  • Bachelor’s degree in Engineering, Information Technology, Computer Science, Data Engineering, or related field, or four (4) years of equivalent practical experience.
  • Six (6) years of relevant experience.
  • Hands-on experience building AWS-native data lake or lakehouse architectures on Amazon S3.
  • Strong experience developing production ETL/ELT pipelines with Python and PySpark, including data modeling, transformation, and performance tuning.
  • Hands-on experience with Apache Iceberg, including ACID transactions, schema evolution, time travel, and query optimization.
  • Advanced SQL skills supporting analytical workloads, reporting, and data visualization.
  • Proven experience with data governance, cataloging, lineage, and access control using AWS services.
  • Knowledge of AWS security fundamentals: IAM, KMS, secrets management, network security, logging, SDLC.
  • Proven experience with Infrastructure as Code (IaC) and operating data platforms across environments.
  • Experience with CI/CD pipelines for data workflows with testing, deployment, environment promotion, and rollback.
  • Troubleshooting distributed data workloads, performance optimization, and cost management skills.
  • Excellent collaboration and communication skills to coordinate with cross-team stakeholders.

Would Be Nice to Have

  • Experience with Databricks, Delta Lake, migrating workloads to AWS-native services, and Apache Iceberg.
  • Familiarity with AWS Step Functions, MWAA, Kinesis, DMS, Lambda, MSK, or similar services.
  • Experience with modern DevOps tools: Git, Terraform, CloudFormation, Jenkins, CodePipeline, GitHub Actions, Docker.
  • Knowledge of BI and visualization tools like Amazon QuickSight, Tableau, Power BI.
  • Familiarity with AI-assisted coding tools such as GitHub Copilot, ChatGPT, Cursor, or Kiro.
  • Knowledge of graph modeling, ontology, taxonomy, entity resolution, and hybrid retrieval techniques.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · World Congress 2024

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:43 min

AWS infrastructure stack and data flow pipeline overview

Artem Volk Artem Volk +1 · World Congress 2024

51 sec

Assessing GPT-4o performance for pull request feedback

Merrill Lutsky Merrill Lutsky · World Congress 2025

Videos

See all

Related articles

See all