Title: Azure Data Lead

StratG Inc
New York, NY, United States
2 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Microsoft Azure Cloud Engineering Information Engineering Extract Transform Load (ETL) Data Systems Software Design Patterns Distributed Systems Apache Hive Python (Programming Language) Microsoft SQL Server Object-Oriented Software Development Oracle (Applications)
+16 more
Oracle SQL Developer Performance Tuning Standard Sql Azure Data Lake Software Engineering SQL Databases Data Processing Azure Data Factory Apache Spark Fastapi Data Lakes Pyspark Data Lakehouse Azure Synapse Analytics Data Pipelines Databricks

Job description

The Azure Data Lead will be responsible for leading the modernization and migration of existing Python object-oriented applications into scalable PySpark and Spark SQL data-processing solutions on Azure Databricks. This role requires a strong blend of software engineering, data engineering, cloud architecture, and performance optimization., * Analyze existing Python OOP applications and redesign single-node processing logic for distributed Spark execution.

  • Design, develop, and deploy enterprise-scale data pipelines on Azure Databricks; build reusable PySpark frameworks and utility modules.
  • Implement Delta Lake solutions using the Bronze-Silver-Gold architecture.
  • Build robust ETL/ELT pipelines with Azure Data Factory, ADLS Gen2, and Azure Synapse Analytics.
  • Implement data quality, reconciliation, validation, and monitoring frameworks.
  • Optimize Spark jobs (partitioning, bucketing, caching, broadcast joins, Adaptive Query Execution, Delta optimization) and benchmark converted applications against original Python implementations.

Requirements

Core Skills:

  • Python (expert), OOP, and advanced Python design patterns
  • PySpark, Spark SQL, and SQL
  • Azure Databricks, Azure Data Factory, ADLS Gen2
  • Apache Spark, Delta Lake, Data Lakehouse architecture, distributed computing

Must-have: Python, Azure Databricks, Azure Data Factory (ADF), MS SQL, Oracle PL/SQL.

Good to have: PySpark; certifications in Azure Data Factory, Azure Databricks, SQL, Oracle, or Python.

Experience & Expected Outcome

The ideal candidate for this role will be a senior data engineering leader with proven delivery of large-scale Databricks modernization programs. The expected outcome is to have existing Python applications converted into scalable, cost-efficient, enterprise-grade data solutions on Azure Databricks with proven performance parity.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:33 min

Connecting frontends via a FastAPI proxy backend layer

Saoussen Chaabnia Saoussen Chaabnia · Europe 2026 Virtual

2:27 min

Managing traffic and tracking costs with Databricks Unity Catalog

Viktoria Semaan Viktoria Semaan · World Congress 2026 Europe

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:50 min

Executing LoRA fine-tuning using serverless Databricks AI runtimes

Viktoria Semaan Viktoria Semaan · World Congress 2026 Europe

1:42 min

Introduction to the fast API web framework

Sebastián Ramírez · World Congress 2022

Videos

See all

Related articles

See all