Data Solutions Architect - Azure Databricks

MINDBRIDGE SOLUTIONS INC
United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$187,200.0 - $208,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Application Integration Architecture JIRA Audit Trail Microsoft Azure Software Quality Code Review Continuous Integration Data Architecture Information Engineering Data Governance Data Systems
+41 more
Data Warehousing Software Design Documents Apache Hive Java Database Connectivity JSON Python (Programming Language) Key Management Metadata Microsoft SQL Server Network Architecture Open Database Connectivity Scrum Methodology Role-Based Access Control Power BI Azure Active Directory Azure Data Lake SQL Databases Data Streaming Systems Integration Tableau (Software) Web Services Extensible Markup Language (XML) File Transfer Protocol (FTP) Data Classification Okta Azure Data Factory Snowflake Apache Spark Microsoft Fabric Data Lakes Pyspark Information Technology Data Lineage Collibra Production Code Bicep Enterprise Integration Data Management Terraform Pagination Databricks

Job description

The Data Solutions Architect is a deeply hands-on technical leader responsible for designing and delivering the end-to-end Azure Databricks platform architecture for a critical Snowflake replatforming initiative in the legal industry. This is not a purely advisory role: the architect is expected to spend approximately 25% of their time writing production code, building reference implementations, and conducting peer code reviews alongside the engineering team. Working as a peer to the Technical Lead, this role owns the architectural vision: from metadata-driven ingestion pipeline patterns through Unity Catalog governance, BI connectivity, and enterprise integrations: and translates that vision into detailed, executable architecture diagrams and standards that guide a team of 8+ Senior Data Engineers. This is a fully remote, independent contractor engagement billed at an hourly rate.

Engagement Details

Role Title Data Solutions Architect - Azure Databricks Engagement Type Independent Contractor - Hourly Rate Work Location Remote Industry Legal Team Context Peer to Technical Lead; guides 8+ Senior Data Engineers Reports To Senior Engineering Manager Primary Initiative Snowflake-to-Azure Databricks Replatforming

Platform Integrations Unity Catalog Collibra Azure DevOps Azure Key Vault Okta / AAD Power BI Tableau ADF ADLS Gen2 Azure Monitor SQL Server Databricks, Solution Architecture & Design

  • Own end-to-end Azure Databricks platform architecture: from raw ingestion through Medallion Lakehouse layers
  • (Bronze/Silver/Gold) to BI consumption: producing detailed, fully documented architecture diagrams that addressboth functional and non-functional requirements (NFRs: scalability, availability, security, performance,maintainability).
  • Design the metadata-driven pipeline framework governing all ingestion patterns: on-premises SQL Server (JDBC,

incremental/CDC), REST/SOAP API integrations (auth, pagination, rate-limiting, retry), and flat file ingestion (CSV,JSON, XML via ADLS Gen2/SFTP landing zones).

  • Architect the migration path from legacy Snowflake and Azure Data Factory (ADF) to native Databricks tooling:
  • including Databricks Workflows, Databricks Autoloader, and Delta Live Tables: minimizing replatforming risk whilemaximizing platform capability adoption.
  • Define metadata store design (pipeline configuration tables, ingestion control frameworks) that enables engineers

to onboard new data sources without bespoke pipeline code.

  • Establish Unity Catalog architecture: workspace federation, catalog/schema/table hierarchy, data classification,

column-level security, row filters, and audit log strategy.synchronization with Unity Catalog.SCIM provisioning into Databricks and Unity Catalog.BI & Consumption Layer ArchitectureSQL connector, DirectQuery, and Fabric integration) and Tableau (via Databricks JDBC/ODBC): includingsemantic layer patterns, aggregation strategies, and performance optimization.workloads without degrading pipeline performance.legal business users.Engineering Enablement & Technical Guidanceimplementation guides, runbooks, and reference implementations that 8+ Senior Data Engineers can executeconfidently.approved patterns early and guide corrective action.DevOps: ensuring architecture standards are enforced through automation where possible.implementations, metadata framework components, and reusable pipeline modules in PySpark and Python;reviewing engineer pull requests for architectural conformance, code quality, and performance.legacy ADF patterns; provide clear migration rationale and transition guidance to the team.Security, Governance & Complianceensure no plaintext credentials in code or notebooks.classification, and privilege-protection controls aligned with legal industry requirements.prevention for the Databricks workspace.Replatforming & Stakeholder Engagementtransformation logic to inform migration sequencing and effort estimation.decommission milestones, risk mitigation, and new capability delivery.decisions against functional and NFR requirements; translate business constraints into platform design guardrails.integration architecture diagrams, and ADRs (Architecture Decision Records).Education and/or Experience

  • Design Collibra integration patterns for bidirectional lineage, data stewardship workflows, and business glossary
  • Architect Okta/Azure Active Directory integration for platform authentication, service principal management, and
  • Design the connectivity architecture between the Databricks Lakehouse and BI tools: Power BI (via Databricks
  • Define Databricks SQL Warehouse sizing, clustering policies, and access control patterns to support concurrent BI
  • Establish certified dataset and Gold layer table design standards that support reliable, governed BI consumption by
  • Serve as the technical authority for the engineering team: translate architecture decisions into detailed
  • Conduct architecture reviews and design walkthroughs at key delivery milestones; identify deviations from
  • Partner with the Technical Lead on code review standards, branching strategy, and CI/CD pipeline design in Azure
  • Dedicate approximately 25% of working time to hands-on coding and peer code review: writing reference
  • Champion Databricks-native tooling adoption (Autoloader, DLT, Databricks Asset Bundles, Unity Catalog) over
  • Design secrets management and credential rotation patterns using Azure Key Vault integration within Databricks;
  • Define data governance standards including data lineage (Unity Catalog + Collibra), data quality SLAs, PII
  • Establish network architecture standards: private endpoints, VNet injection, IP access lists, and data exfiltration
  • Lead technical discovery on the existing Snowflake and ADF estate: catalog existing pipelines, data models, and
  • Define the phased replatforming roadmap in collaboration with the Senior Engineering Manager: balancing legacy
  • Engage directly with legal business stakeholders, compliance teams, and BI consumers to validate architecture
  • Produce and maintain living architecture documentation: solution design documents (SDDs), data flow diagrams

Requirements

Do you have experience in XML?, * Bachelor’s degree in Computer Science, Engineering, Data Science, or a related field.

  • 10+ years of data engineering and/or data architecture experience, including hands-on delivery of large-scale cloud
  • data platforms.
  • 5+ years of deep, production-level Databricks experience: including Unity Catalog, Delta Lake, Databricks

Workflows, Autoloader, Databricks SQL, and cluster/workspace administration.Server (JDBC, CDC), REST/SOAP API, and flat file ingestion patterns at enterprise scale.replatforming initiative.clustering, optimize/ZORDER, time travel, and incremental processing patterns.ability to design performant semantic layer and Gold table patterns for BI consumption.security, row filters, and audit logging.provisioning for Databricks.Azure networking (VNet injection, private endpoints).diagrams, integration architecture diagrams, and ADRs.teams, and engineering staff simultaneously.Preferredor eDiscovery.pipelines.and Unity Catalog provisioning.architecture documentation and engineering productivity.

  • Proven experience architecting and delivering metadata-driven pipeline frameworks covering on-premises SQL
  • Demonstrated experience leading a Snowflake-to-cloud-native migration or comparable legacy data warehouse
  • Strong command of Medallion/Lakehouse architecture and Delta Lake internals: schema evolution, liquid
  • Hands-on experience integrating Databricks with Power BI (DirectQuery, Fabric) and Tableau (JDBC/ODBC);
  • Experience designing Unity Catalog governance architecture including catalog hierarchy, RBAC, column-level
  • Proficiency in Azure Key Vault integration, service principal management, and Okta/Azure AD SSO and SCIM
  • Advanced proficiency in PySpark, Spark SQL, Python, and SQL; working knowledge of Scala.
  • Strong Azure ecosystem experience: ADLS Gen2, Azure Data Factory, Azure DevOps (CI/CD), Azure Monitor, and
  • Proven ability to produce high-quality architecture deliverables: solution design documents, end-to-end data flow
  • Excellent communication skills with demonstrated ability to engage senior business stakeholders, compliance
  • Databricks Certified Data Engineer Professional or Databricks Certified Associate Developer for Apache Spark.
  • Microsoft Certified: Azure Data Engineer Associate or Azure Solutions Architect Expert.
  • Master’s degree in Computer Science, Engineering, or a related field.
  • Familiarity with legal industry data domains: matter management, billing analytics, contract lifecycle management,
  • Experience with data quality frameworks such as Great Expectations or Databricks DQX within Databricks
  • Hands-on experience with Infrastructure as Code (IaC) using Terraform or Bicep for Azure Databricks workspace
  • Familiarity with Microsoft Fabric and its integration with Databricks and Power BI in a hybrid lakehouse topology.
  • Experience with VS Code for Databricks development workflows; familiarity with Claude Code for AI-assisted
  • Familiarity with Agile/Scrum delivery using Azure DevOps Boards or Jira.

Benefits & conditions

$90 - $100 an hour - Contract

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:33 min

Introduction to security advocacy and automation testing

Chris Heilmann +2 · LIVE

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

4:37 min

Architecting single sign-on flows across multiple application domains

Gift Egwuenu · WWC 2023

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters · WWC 2025

Videos

See all

Related articles

See all