Databricks Data Engineer

Alephys LLC
United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$75,000.0 - $140,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Amazon S3 Microsoft Azure Cloud Computing Cloud Engineering Continuous Integration Data Architecture Information Engineering Data Governance Data Infrastructure Extract Transform Load (ETL) IBM InfoSphere DataStage
+19 more
DevOps Distributed Systems Github Apache Hive Python (Programming Language) Performance Tuning DataOps Data Processing Data Ingestion Apache Spark Git Pyspark Gitlab-ci Semi-structured Data Data Lineage Terraform Data Pipelines Jenkins Databricks

Job description

We are seeking a highly skilled Databricks Data Engineer to design, build, and optimize our next-generation data architecture. In this role, you will be instrumental in modernizing our data infrastructure, transitioning legacy ETL workloads to highly scalable, cloud-native solutions, and ensuring our data assets are secure, discoverable, and performant. You will leverage the latest capabilities of the Databricks platform-from declarative pipelines to advanced governance-to deliver high-quality data products to the business., * Pipeline Engineering: Design, build, and maintain robust, scalable ETL/ELT pipelines using PySpark and Spark SQL to process large volumes of structured and semi-structured data.

  • Declarative Frameworks: Implement and manage Spark Declarative Pipelines (Delta Live Tables) to simplify pipeline development, automate data quality checks, and streamline operations.
  • Legacy Modernization: Lead the migration of complex workloads from legacy enterprise ETL systems into modernized, scalable PySpark architectures.
  • Data Governance & Security: Architect and enforce centralized data governance, access controls, auditing, and data lineage tracking across the organization utilizing Unity Catalog.
  • CI/CD & Automation: Automate the deployment lifecycle of data pipelines, notebooks, and infrastructure using Databricks Asset Bundles (DABs) integrated with enterprise CI/CD workflows.
  • Performance Optimization: Profile, tune, and optimize complex PySpark jobs and Spark SQL queries for maximum performance and cost-efficiency.

Requirements

Do you have experience in Spark?, * Experience: 5+ years of experience in Data Engineering, with a heavy focus on the Databricks Data Intelligence Platform.

  • Core Languages: Expert-level proficiency in Python (PySpark) and advanced Spark SQL.
  • Databricks Ecosystem: Deep hands-on experience with modern Databricks features, specifically Spark Declarative Pipelines and Unity Catalog for centralized governance.
  • DevOps / DataOps: Proven ability to implement CI/CD pipelines for data assets using Declarative Asset Bundles (DABs), Git, and automation tools (e.g., GitHub Actions, Jenkins, or GitLab CI).
  • Architecture & Migration: Strong understanding of distributed computing principles and experience migrating legacy on-premise ETL logic (e.g., DataStage, Informatica) to cloud-native Spark environments.
  • Problem Solving: Strong analytical skills with the ability to troubleshoot complex data processing issues and optimize massive data transformations.

Nice-to-Haves

  • Experience with event-driven data ingestion architectures and storage-layer triggers (e.g., S3).
  • Familiarity with cloud infrastructure (AWS, Azure, or GCP) and infrastructure-as-code (Terraform).
  • Experience building and designing internal frameworks to accelerate data pipeline development.

Benefits & conditions

$75,000 - $140,000 a year - Full-time, Contract

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all