> Markdown version of [/jobs/ext/2634507-data-architect-databricks-lakehouse-migration](https://www.wearedevelopers.com/jobs/ext/2634507-data-architect-databricks-lakehouse-migration). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Architect - Databricks Lakehouse Migration - **Company:** Data Inc - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $104,000.0 - $124,800.0 - **Contract:** Permanent contract - **Skills:** Architectural Patterns, Microsoft Azure, Big Data, Databases, Data Architecture, Extract Transform Load (ETL), Role-Based Access Control, SQL Databases, Tokenization, Data Classification, Data Lakes, Terraform, Databricks - **Published:** August 5, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=7bc69db4a79067b9 ## About the Role * 5+ years of data architecture experience * At least one large-scale data warehouse-to-cloud migration (100+ tables) * Hands-on experience with Azure Databricks and Unity Catalog * Production experience with data tokenization, masking, or format-preserving encryption * Strong knowledge of Delta Lake and medallion architecture patterns * SQL fluency and schema mapping across heterogeneous database platforms * Experience defining data classification schemes and mapping to technical controls Preferred Qualifications * Experience with Protegrity Database Protector (or Vormetric, Voltage, or similar) * Financial services or regulated industry background (GLBA, NCUA, FFIEC) * Experience designing detokenization access models tied to identity/RBAC systems * Familiarity with Purview or similar data catalog/classification tool * Terraform or Databricks Asset Bundles experience * Databricks Certified Data Engineer or Data Architect certification ## Description We are seeking an experienced Data Architect to lead the migration of 158 tables from an on-premise data warehouse into a governed Azure Databricks lakehouse. This is a hands-on architecture-and-delivery role - you will own the target data model design, tokenization boundary, medallion architecture, Unity Catalog governance, and migration execution from planning through validation. This engagement involves sensitive financial data in a regulated environment, requiring data to be tokenized at the source before it leaves the on-premise environment. Responsibilities * Design the target Databricks/Unity Catalog architecture: catalog/schema structure, medallion zone mapping (bronze/silver/gold), and table-level ownership * Classify source tables and columns for sensitivity (PII, NPI, account numbers) and map each to a tokenization policy * Design the tokenization boundary ensuring sensitive fields are tokenized on-premise before any data movement to Azure * Produce migration sequencing plan with wave grouping, dependency ordering, and cutover criteria * Define source-to-target data model including schema transformation, type mapping, and reconciliation approach * Design detokenization access model: authorized roles, service identities, and conditions mapped to Unity Catalog grants * Establish data quality and validation gates between source and target for each migration wave * Produce compliance documentation and as-built architecture runbooks ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Kubernetes and Microservices with Multi-Model Databases](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases](https://www.wearedevelopers.com/videos/1146-fault-tolerance-and-consistency-at-scale-harnessing-the-power-of-distributed-sql-databases) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Why Event-Driven Architecture Isn’t About Speed (and When You Actually Need It)](https://www.wearedevelopers.com/magazine/745-why-event-driven-architecture-isn-t-about-speed-and-when-you-actually-need-it)