Databricks Data Lake Engineer

AIT Global, Inc.
Jersey City, NJ, United States
about 2 months ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Airflow Amazon Web Services Amazon S3 Big Data Cloud Database Cloud Engineering Continuous Integration Information Engineering Extract Transform Load (ETL) Data Migration Data Warehousing
+22 more
DevOps Distributed Computing Environment Apache Hadoop Python (Programming Language) Open Database Connectivity Cloud Services Azure Machine Learning SQL Databases Feature Engineering Data Ingestion Azure Data Factory Informatica Powercenter Apache Spark Change Data Capture Git Data Lakes Data Lineage Apache Kafka Machine Learning Operations Restful APIs Legacy Systems Databricks

Job description

The Databricks Data Lake Engineer will design, build, and optimize large scale data pipelines across the full lifecycle of data ingestion, migration, curation, and consumption within a modern lakehouse architecture. This role requires hands on expertise with Databricks, Delta Lake, Spark, and cloud-native data platforms, along with the ability to collaborate effectively with business stakeholders, client, architects, and AI/ML teams., * Data Ingestion - Build scalable ingestion pipelines using Spark, Autoloader, LakeFlow, Informatica, Delta Live Tables, and cloud-native connectors (Kafka, REST, ODBC, CDC - Change Data Capture).

  • Data Migration - Lead migration of legacy data warehouses, Hadoop clusters, or on prem systems into S3/Delta Lake.
  • Data Curation - Implement bronze silver gold architecture, enforce quality checks, schema evolution, and governance.
  • Data Consumption - Deliver curated datasets for BI, analytics, dashboards, and downstream applications.
  • AI/ML Enablement - Partner with data scientists to prepare feature stores, optimize ML-ready datasets, and support model deployment workflows.
  • Develop and maintain CI/CD pipelines for Databricks jobs, notebooks, and workflows.
  • Optimize Spark/SQL/Python jobs for performance, cost efficiency, and reliability.
  • Implement security, governance, and compliance using Unity Catalog, data lineage, and access controls.
  • Collaborate with cross-functional teams and communicate technical concepts clearly to non-technical stakeholders.

Requirements

  • 16 years of education with minimum 5+ years of hands-on experience in data engineering with cloud platforms (AWS preferred).
  • Strong expertise in Databricks, Delta Lake, Apache Spark, and distributed data processing.
  • Experience with Python, SQL, and ETL/ELT frameworks.
  • Proven experience with data migration from legacy systems to cloud data lakes.
  • Deep understanding of data modeling, curation layers, and consumption patterns.
  • Familiarity with ML workflows, feature engineering, and model operationalization.
  • Experience with DevOps, Git, CI/CD, and job orchestration tools.
  • Excellent communication skills with the ability to translate complex concepts into clear business language.

Preferred Qualifications:

  • Experience with AWS Databricks, Azure Data Factory Glue, Airflow, Kafka, Informatica, or similar ingestion tools.
  • Knowledge of Unity Catalog, Delta Sharing, and enterprise governance frameworks.
  • Exposure to AI/ML platforms, MLOps, or Databricks Feature Store.
  • Certifications: Databricks Data Engineer Associate/Professional, AWS and Azure Data Engineer.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all