WeAreDevelopers LIVE β€’ Dec 7, 2021

Fully Orchestrating Databricks from Airflow

Alan Mazankiewicz

Native Databricks scheduling often falls short for complex enterprise pipelines. Discover how integrating Apache Airflow unlocks dynamic, code-driven DAGs for highly scalable data orchestration.

Pause
Mute Enter Fullscreen
#1 about 1 min

Introduction to the speaker and IT solutions company

A machine learning engineer introduces their background and the IT solutions company Ovex.

#2 about 3 min

Overview of Databricks and interactive data processing

Managed Spark clusters enable interactive data analysis using Jupyter notebooks and parallel processing.

#3 about 3 min

Scheduling workflows with Databricks jobs and tasks

Multi-task job orchestration allows defining dependencies between scheduled notebooks and packages within a workspace.

#4 about 3 min

Introduction to Apache Airflow for advanced orchestration

The Airflow UI provides visibility into task schedules, success rates, and log outputs for complex dependencies.

#5 about 3 min

Writing dynamic Airflow DAGs with Python code

Python operators enable dynamic task generation and conditional branching based on external variables or execution dates.

#6 about 6 min

Executing Databricks jobs with built-in Airflow operators

The Databricks Run Now and Submit Run operators trigger predefined or dynamically configured Spark jobs via API calls.

#7 about 9 min

Building generic custom operators for Databricks APIs

Subclassing the base Airflow operator allows interacting with unsupported Databricks endpoints like file systems and custom clusters.

#8 about 2 min

Deploying and running Airflow in cloud environments

While local testing is possible, production Airflow typically runs on Kubernetes or managed services like Cloud Composer.

#9 about 3 min

Optimizing Spark performance for large data volumes

Analyzing the Spark UI helps identify bottlenecks and optimize processing times independently of the Airflow scheduler.

#10 about 1 min

Comparing Airflow orchestration with Azure Data Factory

Data Factory offers a UI-driven alternative for scheduling Spark workloads compared to code-first Airflow deployments.

#11 about 2 min

Selecting the ideal data stack for distinct workloads

Exploratory work benefits from interactive notebooks and Pandas, while production data products rely on Airflow and Databricks.

#12 about 2 min

Managing worker nodes and compute quotas in Databricks

Cluster sizes are primarily limited by cloud provider quotas rather than inherent Databricks constraints.

#13 about 3 min

Enhancing Databricks tooling for software engineering workflows

Future platform improvements should bridge the gap between data science notebook environments and standard application deployment practices.

#14 about 3 min

Learning Python and navigating a practical programming career

A logical approach to problem-solving helps beginners transition from theoretical statistics to practical software development.

#15 about 3 min

Daily routines and software engineering team collaboration

Effective engineering roles combine focused coding time with continuous communication and code reviews across teams.

Matching moments

1:32 min

Handling untyped ingestion and comparing dbt against Spark

Matthias Niehoff Matthias Niehoff Β· World Congress 2023

1:43 min

AWS infrastructure stack and data flow pipeline overview

Artem Volk Artem Volk +1 Β· World Congress 2024

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon Β· World Congress 2026 Europe

5:19 min

Executing queries and scheduling pipeline jobs within DataWorks

Qiyang Duan Β· LIVE

3:33 min

Refactoring data science workflows using Rapids QDF and Pandas

Paul Graham Paul Graham Β· LIVE

2:38 min

Solving complex pipeline orchestration with Argo Workflows

Hauke Brammer Β· World Congress 2023