Databricks Architect
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+43 more
Job description
Data Pipeline Development: Design, code, and deploy robust and scalable batch and streaming data pipelines using PySpark, Spark SQL, and Lakeflow (Lakeflow Connect for ingestion and Lakeflow Declarative Pipelines) to ingest data from sources such as Point-of-Sale (POS), e-commerce platforms, loyalty systems, and marketing clouds. Data Modeling & Transformation: Implement complex data transformations and business logic within the Medallion architecture (Bronze, Silver, Gold layers). Build and optimize the final “Gold” dimension tables that will serve as the single source of truth. Define governed business KPIs on top of the Gold layer using Unity Catalog Metric Views (business semantics) so metrics are computed once and reused consistently across BI, SQL, and Genie. Data Quality: Implement data quality frameworks and cleansing routines to ensure the accuracy and trustworthiness of the Customer 360 data. Performance Optimization: Proactively monitor, debug, and tune Databricks jobs and Spark clusters for performance and cost-efficiency. Implement best practices for partitioning, caching, and data layout in Delta Lake. Infrastructure as Code (IaC) & CI/CD: Work with DevOps teams to manage Databricks environments, clusters, and job deployments using tools like Terraform and AWS DevOps/GitHub Actions. Champion and implement CI/CD best practices for data pipelines. Data Governance & Security: Implement data governance features within Databricks Unity Catalog, including data lineage tracking, fine-grained access controls (row/column-level security and ABAC policies), tags, and data masking to ensure compliance and security across BI and AI/agent workloads. C ollaboration: Partner closely with Functional Consultants, Data Scientists, and Analytics Engineers to understand their data requirements and deliver well-structured, consumption-ready datasets. Semantic Modeling & Business Semantics: Build and govern Unity Catalog Business Semantics, authoring Metric Views (measures, dimensions, joins, synonyms, materialization), Domains, Pages, and Glossary terms, and certifying trusted assets so every dashboard, SQL query, notebook, and AI agent works from the same governed definitions. Genie & Agentic Analytics Enablement: Configure and curate AI/BI Genie Spaces and the Genie Ontology (instructions, trusted assets, example queries, synonyms) to deliver accurate natural-language, conversational analytics for business users, and enable consumption through AI/BI Dashboards and Databricks One.
Requirements
Experience: 10+ years of hands-on data engineering experience, with at least 3 years focused on the Databricks/Spark Ecosystem Databricks Expertise: Deep, hands-on expertise with the Databricks Lakehouse Platform, including Delta Lake, Structured Streaming, Lakeflow Declarative Pipelines (formerly Delta Live Tables), Databricks SQL, and cluster/serverless configuration and optimization. Business Semantics & Governance: Hands-on experience with Unity Catalog Business Semantics, including Metric Views (measures, dimensions, materialization), Domains, and Pages/Glossary to define governed, reusable KPIs once and serve them consistently across SQL, BI, and AI agents. Agentic & Conversational Analytics: Working knowledge/exposure of AI/BI Genie (Genie Spaces, Genie Ontology, trusted assets), Genie Code for agentic pipeline/SQL development, Databricks One for business-user consumption, and Mosaic AI for building and serving AI/ML models and agents. Programming Mastery: Expert-level proficiency in Python and PySpark. Advanced SQL skills are essential. Data Warehousing Concepts: Strong understanding of data modeling principles, including dimensional modeling (Kimball), data warehousing concepts, and ETL/ELT design patterns. Cloud Proficiency: Proven experience working with a major cloud provider (Azure, AWS, or GCP), particularly with data storage S3 and related services. Software Engineering Mindset: Experience with software engineering best practices, including version control (Git), code reviews, testing, and CI/CD. Certification: Databricks, Access Control, Amazon Simple Storage Service (S3), Amazon Web Services (AWS), Apache Spark, Artificial Intelligence (AI), Artificial Intelligence (AI) Agents, Best Practices, Business Analysis, Business Intelligence, Business Model, Business Writing, Caching, Cisco Unity, Cloud Computing, Code Reviews, Continuous Deployment/Delivery, Continuous Integration, Data Analysis, Data Management, Data Modeling, Data Quality, Data Science, Data Storage, Data Warehousing, Database Extract Transform and Load (ETL), Debugging Skills, Design Patterns Programming Methodologies, DevOps, Dimensional Modeling, GCP (Good Clinical Practices), Git, GitHub, Information/Data Security (InfoSec), Legal, Loan Funding, Maintain Compliance, Marketing, Metrics, Microsoft Windows Azure, Ontology, Performance Metrics, Performance Tuning/Optimization, Python Programming/Scripting Language, Reporting Dashboards, SQL (Structured Query Language), Sales, Security Compliance, Software Engineering, Source Code/Configuration Management (SCM), Student Loans, Team Player, Testing, Training Data Sets, eCommerce
Benefits & conditions
Discretionary Annual Incentive. Comprehensive Medical Coverage: Medical & Health, Dental & Vision, Disability Planning & Insurance, Pet Insurance Plans. Family Support: Maternal & Parental Leaves. Insurance Options: Aut& Home Insurance, Identity Theft Protection. Convenience & Professional Growth: Commuter Benefits & Certification & Training Reimbursement. Time Off: Vacation, Time Off, Sick Leave & Holidays. Legal & Financial Assistance: Legal Assistance, 401K Plan, Performance Bonus, College Fund, Student Loan Refinancing.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Top Big Data Technologies That You Need to Know
Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?
Dev Digest 132 - Binging WADFlix?
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production