Databricks Data Architect

ITRANSITION, INC.
United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Unity 3d Artificial Intelligence Amazon Web Services Computing Platforms Audit Trail Microsoft Azure Cloud Computing Data Architecture Data Validation Data Cleansing Data Deduplication Data Governance
+14 more
Extract Transform Load (ETL) Graph Database Role-Based Access Control Search Technologies SQL Databases Workflow Management Systems Large Language Models Apache Spark Data Lakes Pyspark Data Management Data Pipelines User Identification Databricks

Job description

Our client is building a unified Master Data Management platform that consolidates data from many applications into a single source of truth. On top of this platform an AI layer is being built - either as a separate layer over the unified data, or embedded directly into the ETL processes to deliver clean, connected, and enriched data. The project domain is Occupational Health & Safety, Incident Management, Risk Management, and global regulatory frameworks. The platform must provide trustworthy, connected, and explainable data for incident management, risk assessment, and compliance with global regulatory requirements. This is a hands-on architecture role. The person will own the platform design and also build reference pipelines and data models, additionally define AI/RAG patterns together with the engineering team. Tech stack: Databricks (Spark, PySpark, SQL, Delta Lake, Unity Catalog, Lakeflow), ETL/ELT, knowledge graphs / GraphFrames, semantic layer (Unity Catalog metric views), AI Search / Vector Search, RAG / LLM tooling, cloud infrastructure (AWS / Azure / GCP). office remotePoland, * Own the end-to-end architecture of a Databricks-based MDM platform for occupational health, safety, incident, risk, and regulatory data

  • Design ingestion and transformation patterns using Databricks, Spark, PySpark, SQL, Delta Lake, Unity Catalog, and Lakeflow where appropriate
  • Define canonical data models, golden-record logic, entity-resolution rules, and survivorship strategies across heterogeneous source systems
  • Build a semantic layer that provides consistent definitions for incidents, organizations, locations, hazards, controls, risks, regulations, corrective actions, and compliance metrics
  • Design graph-based relationship models for linking entities across systems and enriching downstream analytics and AI use cases
  • Architect AI/RAG capabilities for semantic search, regulatory lookup, incident enrichment, data validation, and source-grounded answers over governed enterprise data
  • Embed data quality, lineage, governance, access control, auditability, and monitoring into the platform from the start
  • Partner with product, engineering, compliance, and analytics teams to convert domain requirements into scalable architecture and implementation patterns

Requirements

  • Strong production experience with Databricks Lakehouse architecture, including Spark, PySpark, SQL, Delta Lake, Unity Catalog, and workflow orchestration
  • Hands-on experience designing and building ETL/ELT pipelines for batch and incremental ingestion, cleansing, normalization, deduplication, and enrichment
  • Practical experience with MDM: golden records, survivorship/merge rules, trust ranking, identity resolution, duplicate detection, SCD, and exception workflows
  • Strong data modeling skills for analytical, operational, and semantic consumption patterns
  • Experience designing a semantic layer with shared business definitions, governed metrics, reusable dimensions, and consistent entity definitions
  • Experience with data quality and observability: pipeline SLAs, schema drift, CDC, data contracts, dead-letter handling, and source-to-master reconciliation
  • Experience implementing data governance and security: Unity Catalog lineage, RBAC/ABAC, row/column-level security, PII handling, and regulatory traceability
  • Ability to translate business requirements from product, compliance, and engineering stakeholders into scalable data architecture, * Experience with Databricks Lakeflow Connect, Lakeflow Spark Declarative Pipelines, and Lakeflow Jobs
  • Experience with Unity Catalog metric views or comparable semantic-layer technologies
  • Experience with knowledge graphs, graph analytics (e.g. GraphFrames), or graph-based entity resolution - linking people, organizations, locations, incidents, hazards, controls, regulations, assets, and corrective actions
  • Experience building AI/RAG solutions over enterprise data using AI Search / Vector Search, embeddings, metadata filtering, retrieval evaluation, and source-grounded generation with citations
  • Experience with ML-based data enrichment, classification, anomaly detection, or entity matching
  • Experience in regulated domains such as occupational health and safety, incident management, risk, compliance, ESG, insurance, healthcare, or industrial operations

Benefits & conditions

  • Projects for such clients as PayPal, Wargaming, Xerox, Philips, Adidas and Toyota
  • Competitive compensation that depends on your qualification and skills
  • Career development system with clear skill qualifications
  • Flexible working hours aligned to your schedule
  • Options to work remotely
  • Corporate medical insurance covering services of private and public medical centers
  • English courses online
  • Corporate parties and events for employees and their children
  • Internal conferences, workshops and meetups for learning and experience sharing
  • Gym membership compensation
  • 5 days of paid sick leave per year with no obligation to submit a sick-leave certificate

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on itransition.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:27 min

Managing traffic and tracking costs with Databricks Unity Catalog

Viktoria Semaan Viktoria Semaan · WWC Europe 2026

2:36 min

Development tools for spatial computing and drones

Zaid Zaim Zaid Zaim · WWC 2023

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:59 min

Key takeaways and accessing the Databricks developer toolkit

Viktoria Semaan Viktoria Semaan · WWC Europe 2026

2:50 min

Executing LoRA fine-tuning using serverless Databricks AI runtimes

Viktoria Semaan Viktoria Semaan · WWC Europe 2026

Videos

See all

Related articles

See all