Lead Data Engineer

Smart Folks Inc
Austin, TX, United States
about 1 month ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Airflow Amazon Web Services Amazon S3 Data Analysis Apache HTTP Server Microsoft Azure Big Data BigQuery Cloud Computing Continuous Integration Information Engineering
+24 more
Data Governance Data Transformation Database Queries DevOps Distributed Computing Environment Distributed Data Store Apache Hive Performance Tuning Query Optimization Cloudera Simple Data Format Workflow Management Systems Parquet Google Cloud Git Data Lakes Pyspark Infrastructure Automation Frameworks Avro Data Management Presto Data Lakehouse Data Pipelines Databricks

Job description

We are seeking a highly skilled and strategic Lead Data Engineer with strong expertise in PySpark Apache Iceberg, Trino, and modern Data Lakehouse architectures. The ideal candidate will be responsible for designing and driving enterprise-scale data platforms that enable analytics, AI/ML, and business intelligence across global organizations., * Lead the architecture, design, and implementation of large-scale distributed data platforms.

  • Build and optimize high-performance data pipelines using PySpark and distributed computing frameworks.
  • Design and manage Data Lakehouse solutions using Apache Iceberg for schema evolution, time travel, partition optimization, and data governance.
  • Architect and optimize federated query solutions using Trino across multiple data sources.
  • Drive enterprise data migration and modernization initiatives from traditional warehouses to Lakehouse architectures.
  • Partner with business stakeholders, product owners, architects, and customer teams to translate business requirements into scalable technical solutions.
  • Establish best practices for data modelling, performance tuning, security, governance, and observability.
  • Mentor a team of data engineers and provide technical leadership across delivery streams.
  • Evaluate and incorporate emerging technologies in Data Engineering, Analytics, and AI.
  • Support pre-sales discussions, solution proposals, estimations, and customer presentations.

Requirements

This is a strategic customer-facing role requiring strong technical leadership, architecture expertise, stakeholder management, and the ability to influence data transformation initiatives., * 10+ years of experience in Data Engineering and Big Data ecosystems.

  • Expert knowledge of PySpark and Spark SQL.
  • Strong hands-on experience with Apache Iceberg.
  • Strong experience with Trino (Presto) query engine.
  • Experience building large-scale batch and near-real-time pipelines.
  • Strong SQL skills and query optimization expertise.
  • Experience with Data Lake technologies and cloud-based analytics platforms.
  • Knowledge of data modelling and distributed storage concepts.
  • Experience with orchestration tools such as Airflow or equivalent.
  • Experience working with file formats such as Parquet, ORC, and Avro.
  • Exposure to CI/CD, Git, DevOps, and Infrastructure as Code practices.

Experience in one or more cloud platforms:

  • AWS (EMR, Glue, S3, Athena, Lake Formation)
  • Azure (Databricks, Data Factory, ADLS)
  • Google Cloud Platform (Dataproc, BigQuery, GCS)

Leadership & Strategic Expectations

  • Ability to engage with senior customer stakeholders.
  • Drive technical roadmaps and platform modernization strategies.
  • Lead architecture reviews and governance forums.
  • Identify opportunities for automation, optimization, and AI-driven solutions.
  • Strong communication and presentation skills.
  • Ability to influence decisions across engineering, product, and business teams.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:02 min

Audience Q&A on data formats and engine tradeoffs

Matthias Niehoff Matthias Niehoff · WWC Europe 2026

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

1:59 min

Evolving roles in AI driven software teams

Ignacio Riesgo Ignacio Riesgo +1 · WWC 2024

1:52 min

Customizing block storage tiers and formats

Ricardo Sueiras Sueiras · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all