DataBricks Data Engineer

Prodapt Corp
Irving, TX, United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Microsoft Azure Code Review Computer Programming Continuous Integration Data Architecture Information Engineering Extract Transform Load (ETL) DevOps Apache Hive Python (Programming Language) Key Management
+23 more
NoSQL Object-Oriented Software Development Performance Tuning Azure Data Lake SQL Databases Systems Integration Cloud Platform System Azure Data Factory ReactJS Snowflake Apache Spark Caching Git Build Management Data Lakes Pyspark Information Technology Deployment Automation Terraform Azure Synapse Analytics Software Version Control Data Pipelines Databricks

Job description

Design, develop, and implement scalable data pipelines using Azure Databricks. Build and optimize ETL/ELT pipelines using PySpark, Python, and SQL. Develop data solutions using the Medallion Architecture (Bronze, Silver, and Gold layers). Work with Delta Lake, Delta tables, and advanced Databricks optimization techniques. Integrate Databricks with Azure services such as ADLS Gen2, Azure Data Factory, Azure Synapse, and Azure Key Vault. Develop and manage Databricks Workflows, Jobs, and Notebooks. Implement data quality, monitoring, error handling, and performance optimization. Collaborate with Data Architects, Data Engineers, Data Scientists, and business stakeholders. Establish best practices for CI/CD, Git/version control, code reviews, and automated deployments. Lead technical discussions and provide guidance to junior engineers. Analyze existing Python OOP applications and redesign single-node processing logic for distributed Spark execution. Design, develop, and deploy enterprise-scale data pipelines on Azure Databricks; build reusable PySpark frameworks and utility modules. Implement Delta Lake solutions using the Bronze, Silver, Gold architecture. Build robust ETL/ELT pipelines with Azure Data Factory, ADLS Gen2, and Azure Synapse Analytics. Implement data quality, reconciliation, validation, and monitoring frameworks. Optimize Spark jobs using partitioning, bucketing, caching, broadcast joins, Adaptive Query Execution, and Delta optimization.

Requirements

10+ years of overall Data Engineering experience, including designing and implementing ETL/ELT pipelines with Azure Data Factory and other Azure services.

Bachelor’s degree in Computer Science, Engineering, or a related field; OR equivalent combination of education and relevant experience. Strong hands-on experience with Azure Databricks; 4+ years of experience across Azure services and Databricks (ADLS, ADF, Azure DevOps, etc.). 7+ years of Python development experience, with the ability to design and build reusable libraries. 4+ years of experience with Snowflake or SQL (No-SQL experience is a plus). Expert-level knowledge of PySpark and Spark SQL. Strong programming experience in Python and SQL. Experience with Delta Lake and Lakehouse architecture. Strong experience with Azure Data Lake Storage (ADLS Gen2). Experience with Azure Data Factory (ADF) and other Azure data services. Strong understanding of ETL/ELT, data modeling, and large-scale data pipelines. Experience with performance tuning and optimization in Databricks/Spark. Experience with Git, CI/CD, and DevOps practices.

Preferred Skills Databricks certifications. Experience with Unity Catalog and Databricks governance/security. Experience with Terraform or Infrastructure as Code. Knowledge of Azure DevOps. Experience designing APIs and integrating with React JS within a cloud platform

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:05 min

Enhancing Databricks tooling for software engineering workflows

Alan Mazankiewicz · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

1:22 min

Analyzing differences between mobile and traditional backend DevOps

Mete Baydar Mete Baydar · World Congress 2025

2:38 min

Overview of Databricks and interactive data processing

Alan Mazankiewicz · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all