Databricks Engineer

Purple Drive Technologies LLC
Malvern, PA, United States
about 2 months ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Amazon S3 Data Analysis Batch Processing Continuous Integration Data Architecture Information Engineering Data Governance Extract Transform Load (ETL) Data Security
+36 more
DevOps Distributed Computing Environment Apache Hive Identity and Access Management Python (Programming Language) Key Management Machine Learning Performance Tuning Query Optimization Standard Sql Search Technologies Data Streaming Systems Integration S3 Bucket Data Processing Feature Engineering Large Language Models Apache Spark Generative AI Git Build Management Data Lakes Pyspark Infrastructure Automation Frameworks Data Lineage Deployment Automation AWS Glue Data Analytics Data Management Machine Learning Operations Terraform Software Version Control Data Pipelines Amazon Redshift Databricks Programming Languages

Job description

We are seeking a highly experienced Senior Databricks Engineer with deep expertise in Apache Spark, Databricks Lakehouse Platform, Delta Lake, AWS, and modern data engineering. The ideal candidate will design and build scalable, secure, and high-performance data pipelines while implementing enterprise-grade data governance, CI/CD, and MLOps practices. Experience with Generative AI, LLMs, and Vector Search is highly desirable., * 10-15 years of experience in Data Engineering

  • Expert in:
  • Databricks Lakehouse Platform
  • Apache Spark
  • Spark SQL
  • PySpark
  • DataFrame API
  • Distributed Data Processing
  • Spark Performance Tuning
  • Query Optimization

Delta Lake & Lakehouse Architecture

  • Delta Lake
  • Medallion Architecture (Bronze, Silver, Gold)
  • Delta Live Tables (DLT)
  • ACID Transactions
  • Z-Ordering
  • Data Optimization
  • Time Travel
  • Schema Evolution

Data Pipelines & Orchestration

  • ETL / ELT Development
  • Databricks Workflows
  • Databricks Jobs
  • Delta Live Tables (DLT)
  • Workflow Automation
  • Batch Processing
  • Streaming Data Pipelines
  • Error Handling
  • Retry Mechanisms

Programming Languages

  • Python
  • PySpark
  • SQL
  • Scala (Preferred)

AWS Cloud & Storage

  • Amazon S3
  • Amazon Redshift
  • AWS Glue
  • Amazon Kinesis
  • AWS Step Functions
  • AWS IAM
  • AWS KMS
  • Cross-Account IAM Roles
  • S3 Bucket Policies

Databricks Governance & Security

  • Unity Catalog
  • Data Lineage
  • Row-Level Security
  • Column-Level Security
  • Data Governance
  • Secure Data Sharing
  • Compliance

Compute & Cost Optimization

  • Cluster Policies
  • Instance Profiles
  • Spot Instances
  • Compute Optimization
  • Cost Management

CI/CD & DevOps

  • Git
  • Databricks Git Folders
  • CI/CD Pipelines
  • Version Control
  • Deployment Automation

AI & MLOps

  • MLflow
  • Model Registry
  • Experiment Tracking
  • Feature Engineering
  • Large Language Models (LLMs)
  • Vector Search
  • Generative AI Integrations, * Design, develop, and maintain scalable ETL/ELT pipelines using Databricks, Apache Spark, and Delta Lake
  • Build and optimize batch and streaming data pipelines processing data from multiple enterprise sources
  • Develop high-performance PySpark and Spark SQL solutions, optimizing distributed processing and execution plans
  • Implement Medallion Architecture (Bronze, Silver, Gold) using Delta Lake best practices
  • Design and automate resilient workflows using Databricks Workflows, Jobs, and Delta Live Tables (DLT)
  • Configure Unity Catalog for enterprise data governance, lineage, and fine-grained access controls
  • Integrate Databricks with AWS services including Amazon S3, Redshift, Glue, Kinesis, and Step Functions
  • Implement secure cloud architectures using IAM roles, S3 bucket policies, and AWS KMS customer-managed encryption keys
  • Optimize Databricks compute resources through cluster policies, instance profiles, and Spot Instances
  • Collaborate with Data Scientists, ML Engineers, Analysts, and BI teams to support analytics, machine learning, and Generative AI initiatives
  • Implement CI/CD pipelines, Git-based development workflows, and deployment automation
  • Utilize MLflow for experiment tracking, model management, and MLOps lifecycle support

Requirements

  • Experience with real-time streaming architectures and event-driven data platforms
  • Hands-on experience with Generative AI, LLMs, Vector Databases, and Retrieval-Augmented Generation (RAG)
  • Databricks Certified Data Engineer Professional or Associate
  • AWS Certified Data Analytics or Solutions Architect certification
  • Experience with Terraform or Infrastructure-as-Code (IaC)
  • Experience supporting enterprise-scale cloud data platforms

CONTRACTOR, FULL_TIME, PART_TIME

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

1:41 min

Visualizing the complex developer journey for JVM ecosystems

Bobur Umurzokov · LIVE

Videos

See all

Related articles

See all