Lead Data Engineer

SDH Systems LLC
San Jose, CA, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Microsoft Azure Big Data Cloud Computing Data Architecture Information Engineering Extract Transform Load (ETL) Data Warehousing Machine Learning Performance Tuning SQL Databases Data Streaming
+12 more
Azure Service Bus Feature Engineering Azure Data Factory Snowflake Apache Spark Data Lakes Pyspark Apache Kafka Azure Synapse Analytics Stream Analytics Data Pipelines Databricks

Job description

  • Design, build, and optimize scalable data pipelines using Databricks, Apache Spark, and Azure technologies.
  • Architect data warehousing solutions, ensuring seamless integration with cloud platforms and structured/unstructured data sources.
  • Collaborate with business stakeholders to understand data needs and develop high-performance analytical solutions.
  • Implement ETL/ELT processes leveraging cloud-based technologies such as Azure Data Factory, Snowflake, and Delta Lake.
  • Ensure data quality, governance, and security compliance while managing large datasets efficiently.
  • Drive performance tuning and optimization for data pipelines, ensuring efficiency across systems.
  • Work closely with cross-functional teams to support machine learning and advanced analytics initiatives.
  • Provide technical leadership and mentorship to junior data engineers, fostering a culture of innovation and continuous improvement.
  • Stay updated on emerging data technologies and recommend strategies to enhance existing architectures.

Requirements

We are seeking a Lead Data Engineer with expertise in Databricks and Data Warehousing to drive data architecture, pipeline development, and optimization efforts. The ideal candidate will play a key role in designing scalable solutions, implementing best practices, and leading data initiatives within a dynamic and collaborative environment., * 8+ years of experience in data engineering, big data processing, and cloud-based solutions.

  • Strong expertise in Databricks, Spark (PySpark/SQL), and Delta Lake architecture.
  • Proven experience in designing and managing data warehouses using Snowflake, Azure Synapse, or equivalent technologies.
  • Deep understanding of data modeling, SQL, and performance optimization.
  • Hands-on experience with Azure Data Factory, Event Hubs, and cloud-based ETL processes.
  • Solid knowledge of real-time streaming technologies (Kafka, Azure Stream Analytics, or similar).
  • Familiarity with ML/AI data pipelines and feature engineering best practices.
  • Strong communication and collaboration skills, with experience working in fast-paced, enterprise environments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

1:33 min

Integrating internal APIs and maintaining data sovereignty

Mahran Meißner Mahran Meißner · WWC Europe 2026

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

Videos

See all

Related articles

See all