Data Lakehouse Engineer

Jconnect Infotech Inc
Cherry Hill, United States
10 days ago
Apply on www.thejobnetwork.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Airflow Amazon Web Services Amazon S3 Apache HTTP Server Cloud Storage Data as a Services Data Governance Extract Transform Load (ETL) Distributed Computing Environment Identity and Access Management JSON Python (Programming Language)
+12 more
Metadata Meta-Data Management Metadata Repositories SQL Databases Workflow Management Systems Enterprise Data Management Data Classification Snowflake Apache Spark Semi-structured Data Data Lineage Collibra

Job description

The Position is for a skilled and detail-oriented Data Lakehouse Engineer with a strong background in Enterprise Data Office (EDO) principles to design, build, and maintain next-generation data lakehouse platforms. This role combines technical excellence with data governance and compliance, supporting enterprise-wide data strategies and analytics initiatives.

The Position will play a key role in the following assignments:

Building snowflake external tables against data saved in AWS S3 in Apache Iceberg.

Using existing pipeline to migrate data from one AWS account in S3 to another AWS account where the data is stored in Apache iceberg tables on S3.

Other Responsibilities:

Design and implement scalable data lakehouse architectures using modern open formats such as Apache Iceberg.

Build and manage Snowflake external tables referencing Apache Iceberg tables stored on AWS S3.

Develop and maintain data ingestion, transformation, and curation pipelines, supporting structured and semi-structured data.

Collaborate with EDO, data governance, and business teams to align on standards, lineage, classification, and data quality enforcement.

Implement data catalogs, metadata management, lineage tracking, and enforce role-based access controls.

Ensure data is compliant with privacy and security policies across ingestion, storage, and consumption layers.

Optimize cloud storage, compute usage, and pipeline performance for cost-effective operations

Requirements

Proven experience building Snowflake external tables over Apache Iceberg tables on AWS S3.

Strong understanding of Iceberg table concepts like schema evolution, partitioning, and time travel.

Proficient in Python, SQL and Spark or other distributed data processing frameworks.

Solid understanding of AWS data services including S3, Glue, Lake Formation, and IAM.

Hands-on with data governance tools (e.g., Collibra/ Alation/ Informatica) and EDO-aligned processes.

Experience implementing metadata, lineage, and data classification standards.

Good to Have Skills:

Experience in Insurance Domain

Experience migrating JSON files to Apache Iceberg tables on AWS S3 using configuration-driven rules.

Familiarity with ETL orchestration tools such as Apache Airflow, DBT, or AWS Step Functions.

Prior exposure to MDM, data quality frameworks, or stewardship platforms

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.thejobnetwork.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

2:02 min

Choosing the right data format and catalog engine

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters · World Congress 2025

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

1:03 min

Understanding the complexity and traps of JSON schema

Clemens Vasters Clemens Vasters · World Congress 2025

Videos

See all

Related articles

See all