> Markdown version of [/videos/336-fully-orchestrating-databricks-from-airflow?t=652](https://www.wearedevelopers.com/videos/336-fully-orchestrating-databricks-from-airflow?t=652). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Fully Orchestrating Databricks from Airflow Native Databricks scheduling often falls short for complex enterprise pipelines. Discover how integrating Apache Airflow unlocks dynamic, code-driven DAGs for highly scalable data orchestration. - **Speakers:** Alan Mazankiewicz - **Event:** WeAreDevelopers LIVE - **Published:** December 7, 2021 - **Duration:** 39:14 - **URL:** https://www.wearedevelopers.com/videos/336-fully-orchestrating-databricks-from-airflow ## Summary As data engineering workflows mature, relying solely on Databricks' built-in scheduling often falls short of complex enterprise needs. While Databricks excels as a managed Apache Spark service for large-scale data processing, Apache Airflow provides a much more robust orchestration layer. By integrating Airflow, engineers can define dynamic, code-driven Directed Acyclic Graphs (DAGs) that leverage advanced trigger rules and dynamic task generation, surpassing the native UI-based multi-task job orchestration within Databricks. The integration relies heavily on the apache-airflow-providers-databricks package, utilizing built-in tools like the DatabricksRunNowOperator for pre-configured jobs and the DatabricksSubmitRunOperator for dynamically defined workloads. Since these operators rely on the Databricks API, managing authentication is securely handled via Airflow's connection and secret management system. Furthermore, Airflow acts purely as a lightweight scheduler, meaning the heavy lifting of data processing and Spark performance optimization remains entirely within the Databricks compute environment. For advanced use cases not covered by built-in operators—such as writing directly to the Databricks File System (DBFS) or spinning up a single, shared all-purpose cluster to process numerous small tasks—developers can subclass Airflow's BaseOperator to build custom generic operators. Utilizing the DatabricksHook and Jinja templating for runtime variables, these custom operators can execute arbitrary API calls. Ultimately, while Databricks is historically tailored toward interactive data science via notebooks, pairing it with Airflow bridges the gap to robust, developer-focused software engineering practices, enabling highly scalable and maintainable data products. **Keywords:** apache airflow orchestration, managed apache spark, databricks jobs api, airflow dag generation, custom airflow operators, databricksrunnowoperator, databrickssubmitrunoperator, airflow connection management, databricks cluster provisioning, dbfs api integration, data engineering pipelines, jinja templating in airflow, spark ui performance optimization, azure data factory vs airflow, python data ecosystem, complex trigger rules ## Chapters 1. **Introduction to the speaker and IT solutions company** (00:00) — A machine learning engineer introduces their background and the IT solutions company Ovex. 1. **Overview of Databricks and interactive data processing** (00:52) — Managed Spark clusters enable interactive data analysis using Jupyter notebooks and parallel processing. 1. **Scheduling workflows with Databricks jobs and tasks** (03:31) — Multi-task job orchestration allows defining dependencies between scheduled notebooks and packages within a workspace. 1. **Introduction to Apache Airflow for advanced orchestration** (05:42) — The Airflow UI provides visibility into task schedules, success rates, and log outputs for complex dependencies. 1. **Writing dynamic Airflow DAGs with Python code** (08:02) — Python operators enable dynamic task generation and conditional branching based on external variables or execution dates. 1. **Executing Databricks jobs with built-in Airflow operators** (10:52) — The Databricks Run Now and Submit Run operators trigger predefined or dynamically configured Spark jobs via API calls. 1. **Building generic custom operators for Databricks APIs** (16:06) — Subclassing the base Airflow operator allows interacting with unsupported Databricks endpoints like file systems and custom clusters. 1. **Deploying and running Airflow in cloud environments** (24:34) — While local testing is possible, production Airflow typically runs on Kubernetes or managed services like Cloud Composer. 1. **Optimizing Spark performance for large data volumes** (25:45) — Analyzing the Spark UI helps identify bottlenecks and optimize processing times independently of the Airflow scheduler. 1. **Comparing Airflow orchestration with Azure Data Factory** (28:15) — Data Factory offers a UI-driven alternative for scheduling Spark workloads compared to code-first Airflow deployments. 1. **Selecting the ideal data stack for distinct workloads** (29:13) — Exploratory work benefits from interactive notebooks and Pandas, while production data products rely on Airflow and Databricks. 1. **Managing worker nodes and compute quotas in Databricks** (30:45) — Cluster sizes are primarily limited by cloud provider quotas rather than inherent Databricks constraints. 1. **Enhancing Databricks tooling for software engineering workflows** (31:47) — Future platform improvements should bridge the gap between data science notebook environments and standard application deployment practices. 1. **Learning Python and navigating a practical programming career** (33:52) — A logical approach to problem-solving helps beginners transition from theoretical statistics to practical software development. 1. **Daily routines and software engineering team collaboration** (36:25) — Effective engineering roles combine focused coding time with continuous communication and code reviews across teams. ## Related Moments - [Handling untyped ingestion and comparing dbt against Spark](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) (from "Enjoying SQL data pipelines with dbt") - [AWS infrastructure stack and data flow pipeline overview](https://www.wearedevelopers.com/videos/944-building-the-platform-for-providing-ml-predictions-based-on-real-time-player-activity) (from "Building the platform for providing ML predictions based on real-time player activity") - [Empowering domain teams with an open data platform](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) (from "From Messy Queries to Scalable Systems - How Data Engineering actually works") - [Executing queries and scheduling pipeline jobs within DataWorks](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) (from "Alibaba Big Data and Machine Learning Technology") - [Refactoring data science workflows using Rapids QDF and Pandas](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) (from "Accelerating Python on GPUs") - [Solving complex pipeline orchestration with Argo Workflows](https://www.wearedevelopers.com/videos/825-mlops-on-kubernetes-exploring-argo-workflows) (from "MLOps on Kubernetes: Exploring Argo Workflows") ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Why Event-Driven Architecture Isn’t About Speed (and When You Actually Need It)](https://www.wearedevelopers.com/magazine/745-why-event-driven-architecture-isn-t-about-speed-and-when-you-actually-need-it) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) ## Related Jobs - [Lead Software Engineer - Data Engineering](https://www.wearedevelopers.com/jobs/ext/2000968-lead-software-engineer-data-engineering) at **Dynatrace** - [Data Scientist](https://www.wearedevelopers.com/jobs/ext/2725456-data-scientist) at **Almedia** - [Machine Learning Engineer](https://www.wearedevelopers.com/jobs/ext/2122890-machine-learning-engineer) at **Twilio** - [Tax Innovation - Data Engineer - Senior Manager](https://www.wearedevelopers.com/jobs/ext/2562027-tax-innovation-data-engineer-senior-manager) at **PwC** - [Azure Solutions Architect Expert, Azure Data Engineer Associate](https://www.wearedevelopers.com/jobs/ext/2536426-azure-solutions-architect-expert-azure-data-engineer-associate) at **PwC** - [Data and AI Specialist](https://www.wearedevelopers.com/jobs/ext/2156264-data-and-ai-specialist) at **Twilio**