Databricks Architect

INFT Solutions inc
United States
14 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Amazon S3 Microsoft Azure Cloud Database Cloud Storage Cluster Analysis Code Review Databases Data Architecture Information Engineering Extract Transform Load (ETL)
+32 more
Data Transformation Data Systems Relational Databases Software Design Patterns Key Management Enterprise Messaging Systems Performance Tuning Cloud Services Migration Manager Standard Sql SQL Databases Talend Strategies of Testing Data Logging Transaction Processing (Computing) Data Processing Enterprise Software Applications Cloud Platform System Data Ingestion Azure Data Factory Apache Spark Data Lakes Pyspark Semi-structured Data Information Technology Apache Kafka Cloud Integration Software Coding Restful APIs Data Pipelines Serverless Computing Databricks

Job description

We are looking for an experienced Senior Databricks Architect to lead the architecture and technical execution of a large-scale migration for a leading Retail/Grocery client.

The architect will be responsible for designing the target Databricks Lakehouse architecture, defining migration patterns, guiding the data engineering team, and ensuring successful modernization of existing Talend-based ETL/ELT workloads onto Databricks.

The role requires strong hands-on expertise in Databricks, Apache Spark, SQL, cloud data platforms, data architecture, ETL modernization, and migration strategy, along with a good understanding of Retail/Grocery business processes and data domains.

Key Responsibilities

Architecture & Solution Design

Lead the end-to-end architecture for migration of Talend ETL/ELT workloads to Databricks.

Define the target-state Databricks Lakehouse architecture, including data ingestion, transformation, storage, orchestration, governance, and consumption layers.

Design scalable, secure, highly available, and cost-optimized data solutions.

Define architecture standards, design patterns, coding standards, and best practices for Databricks development.

Establish migration frameworks and reusable patterns for converting Talend jobs into Databricks/Spark-based pipelines.

Evaluate existing Talend jobs and determine appropriate migration strategies, including re-platforming, re-engineering, consolidation, or retirement.

Provide technical leadership for complex data pipelines and integration scenarios.

Talend to Databricks Migration

Analyze existing Talend workflows, jobs, mappings, dependencies, schedules, and data transformations.

Develop migration strategies for batch and incremental data processing workloads.

Translate Talend transformations and business rules into PySpark/SQL/Databricks implementations.

Identify opportunities to simplify and optimize legacy ETL processes during migration.

Define data reconciliation and validation strategies to ensure parity between Talend and Databricks outputs.

Establish migration sequencing based on business criticality, dependencies, complexity, and risk.

Support migration of high-volume and business-critical data pipelines.

Databricks & Data Engineering

Design and implement solutions using Databricks, Apache Spark, Delta Lake, PySpark, and SQL.

Design Medallion Architecture (Bronze, Silver, Gold) and appropriate data processing patterns.

Implement incremental processing, CDC, SCD Type 1/Type 2, data quality, error handling, and audit frameworks.

Optimize Spark workloads, Delta tables, SQL queries, partitioning, clustering, and job execution.

Leverage Delta Live Tables / Lakeflow capabilities, Workflows, Unity Catalog, and other Databricks platform services where appropriate.

Design data pipelines integrating structured and semi-structured data from databases, files, APIs, and enterprise applications.

Establish monitoring, logging, alerting, operational support, and performance-management patterns.

Cloud & Integration

Strong experience with at least one major cloud platform, preferably Azure or AWS.

Design integration between Databricks and cloud storage, databases, messaging platforms, APIs, and enterprise applications.

Experience with technologies such as ADLS/S3, Azure Data Factory, Kafka, REST APIs, relational databases, and cloud-native services is highly desirable.

Define secure connectivity, networking, secrets management, and access-control patterns.

Retail / Grocery Domain

Work closely with business and technology stakeholders to understand Retail/Grocery data requirements.

Experience with retail data domains such as:

o Product / Item

o Store / Location

o Customer / Loyalty

o Sales / Transactions

o Pricing

o Promotions

o Inventory

o Supply Chain

o Vendors / Suppliers

o Orders

o Forecasting

o Merchandising

Understand retail-specific challenges such as high-volume transaction processing, store-level data, product hierarchies, promotions, pricing, inventory movements, and customer analytics.

Translate business requirements into scalable data architecture and technical solutions.

Technical Leadership

Act as the senior technical authority for the Databricks migration program.

Provide technical direction and mentoring to data engineers and developers.

Conduct architecture/design reviews and code reviews.

Troubleshoot complex production and performance issues.

Collaborate with offshore/onshore teams, enterprise architects, infrastructure teams, security teams, and business stakeholders.

Prepare architecture diagrams, technical design documents, migration plans, and implementation standards.

Communicate technical risks, dependencies, and recommendations effectively to client leadership.

Required Technical Skills

Databricks

Strong hands-on experience with Databricks.

Apache Spark / PySpark.

Delta Lake.

Databricks Workflows/Jobs.

Unity Catalog.

Databricks SQL.

Performance tuning and optimization.

Medallion/Lakehouse architecture.

Experience with modern Databricks data engineering capabilities.

Data Engineering

Strong SQL and data modeling skills.

ETL/ELT architecture.

Batch and incremental data processing.

CDC and SCD implementation.

Data quality and reconciliation frameworks.

Requirements

10+ years overall IT experience, with 5+ years in Databricks / Cloud Data Engineering / Data Architecture

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann Chris Heilmann +3 · LIVE

2:05 min

Enhancing Databricks tooling for software engineering workflows

Alan Mazankiewicz · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

2:38 min

Overview of Databricks and interactive data processing

Alan Mazankiewicz · LIVE

3:44 min

Automating storage savings with S3 intelligent tiering

Sébastien Stormacq · World Congress 2021

Videos

See all

Related articles

See all