Architect- Data Engineering

THE JUDGE GROUP, INC.
Tustin, CA, United States
15 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Compensation
$200,000.0 - $220,000.0
Working hours
Regular working hours
Job source

Tech stack

Airflow Amazon Web Services Computing Platforms Microsoft Azure Cloud Computing Cluster Analysis Code Review Information Engineering Data Security Data Vault Modeling Data Warehousing Dimensional Modeling
+17 more
Distributed Systems Github Apache Hive Python (Programming Language) DataOps Data Streaming Enterprise Data Management Data Processing Google Cloud Apache Spark Data Lakes Pyspark Gitlab-ci Apache Kafka Terraform Data Pipelines Databricks

Job description

We are seeking a visionary and hands-on Senior Data Engineering Architect to design, scale, and optimize our enterprise data platform. In this role, you will define the blueprints for our data estate, leveraging a modern stack centered on Databricks, dbt, and Apache Airflow. You will bridge the gap between complex business strategy and technical implementation, ensuring our data pipelines are scalable, resilient, and cost-effective.

Key Responsibilities

Architecture & Platform Design

Design end-to-end lakehouse architectures on Databricks utilizing Delta Lake and Unity Catalog.

Establish robust governance, schema evolution, and fine-grained data security patterns.

Formulate standard frameworks for data modeling (e.g., Kimball dimensional modeling, Data Vault 2.0).

Optimize infrastructure for optimal price-to-performance across batch and streaming workloads.

Data Pipeline & Orchestration Engineering

Architect modular, reusable transformation frameworks using dbt Core/Cloud integrated with Databricks.

Standardize data processing patterns using PySpark, Delta Live Tables (DLT), and Spark SQL.

Build highly observable, dynamic orchestration workflows using Apache Airflow.

Design cross-DAG dependency models, custom providers, and robust error-handling mechanisms.

DataOps & Engineering Excellence

Drive DataOps maturity by implementing CI/CD pipelines via GitHub Actions, GitLab CI, or Azure DevOps.

Deploy infrastructure-as-code patterns using Terraform and Databricks Asset Bundles (DABs).

Embed automated data quality testing directly into the dbt and Airflow lifecycle.

Define service-level indicators (SLIs) and objectives (SLOs) for pipeline uptime and data freshness.

Leadership & Stakeholder Management

Serve as the principal technical authority and escalation point for data engineering teams.

Mentor senior and mid-level data engineers through code reviews and architectural workshops.

Collaborate with product managers, data scientists, and business leaders to solve data gaps.

Requirements

Overall 15+ Years of experience

10+ years of total experience in data engineering, data warehousing, and distributed systems.

4+ years of dedicated experience architecting production environments within the modern data stack.

Technical Proficiencies

Databricks: Advanced mastery of Photon engine, Unity Catalog, Delta Lake optimization (Z-order, Liquid Clustering), and DLT.

dbt: Expert-level proficiency with compl ex macro development, custom materializations, and multi-project dbt mesh architectures.

Airflow: Deep understanding of Airflow scheduling, custom operators, dynamic task mapping, and infrastructure scaling.

Languages: Elite proficiency in Python (PySpark) and advanced SQL.

Cloud Infrastructure: Strong experience with at least one major cloud ecosystem provider: AWS, Azure, or Google Cloud Platform.

Soft Skills

Strong technical communication skills to distill complex infrastructure designs for non-technical stakeholders.

Natural ability to lead by influence and drive cross-functional engineering initiatives.

Preferred Qualifications

Official Databricks certifications (e.g., Databricks Certified Data Engineer Professional or Solutions Architect).

Active contributor to open-source data communities (dbt, Airflow, or Apache Spark).

Solid foundation in streaming data technologies like Apache Kafka or AWS

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

3:05 min

Audience questions on AI agents and pipeline vectorization

Joy Joy · World Congress 2024

Videos

See all

Related articles

See all