> Markdown version of [/jobs/ext/1026917-lead-data-engineer](https://www.wearedevelopers.com/jobs/ext/1026917-lead-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Data Engineer - **Company:** Smart Folks Inc - **Location:** Austin, TX, United States - **Experience:** Expert - **Contract:** Temporary contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Data Analysis, Apache HTTP Server, Microsoft Azure, Big Data, BigQuery, Cloud Computing, Continuous Integration, Information Engineering, Data Governance, Data Transformation, Database Queries, DevOps, Distributed Computing Environment, Distributed Data Store, Apache Hive, Performance Tuning, Query Optimization, Cloudera, Simple Data Format, Workflow Management Systems, Parquet, Google Cloud, Git, Data Lakes, Pyspark, Infrastructure Automation Frameworks, Avro, Data Management, Presto, Data Lakehouse, Data Pipelines, Databricks - **Published:** June 30, 2026 - **Apply:** https://www.dice.com/job-detail/c8b7fd75-e7f8-4b75-9701-20c66e49d73c ## About the Role This is a strategic customer-facing role requiring strong technical leadership, architecture expertise, stakeholder management, and the ability to influence data transformation initiatives., * 10+ years of experience in Data Engineering and Big Data ecosystems. * Expert knowledge of PySpark and Spark SQL. * Strong hands-on experience with Apache Iceberg. * Strong experience with Trino (Presto) query engine. * Experience building large-scale batch and near-real-time pipelines. * Strong SQL skills and query optimization expertise. * Experience with Data Lake technologies and cloud-based analytics platforms. * Knowledge of data modelling and distributed storage concepts. * Experience with orchestration tools such as Airflow or equivalent. * Experience working with file formats such as Parquet, ORC, and Avro. * Exposure to CI/CD, Git, DevOps, and Infrastructure as Code practices. Experience in one or more cloud platforms: * AWS (EMR, Glue, S3, Athena, Lake Formation) * Azure (Databricks, Data Factory, ADLS) * Google Cloud Platform (Dataproc, BigQuery, GCS) Leadership & Strategic Expectations * Ability to engage with senior customer stakeholders. * Drive technical roadmaps and platform modernization strategies. * Lead architecture reviews and governance forums. * Identify opportunities for automation, optimization, and AI-driven solutions. * Strong communication and presentation skills. * Ability to influence decisions across engineering, product, and business teams. ## Description We are seeking a highly skilled and strategic Lead Data Engineer with strong expertise in PySpark Apache Iceberg, Trino, and modern Data Lakehouse architectures. The ideal candidate will be responsible for designing and driving enterprise-scale data platforms that enable analytics, AI/ML, and business intelligence across global organizations., * Lead the architecture, design, and implementation of large-scale distributed data platforms. * Build and optimize high-performance data pipelines using PySpark and distributed computing frameworks. * Design and manage Data Lakehouse solutions using Apache Iceberg for schema evolution, time travel, partition optimization, and data governance. * Architect and optimize federated query solutions using Trino across multiple data sources. * Drive enterprise data migration and modernization initiatives from traditional warehouses to Lakehouse architectures. * Partner with business stakeholders, product owners, architects, and customer teams to translate business requirements into scalable technical solutions. * Establish best practices for data modelling, performance tuning, security, governance, and observability. * Mentor a team of data engineers and provide technical leadership across delivery streams. * Evaluate and incorporate emerging technologies in Data Engineering, Analytics, and AI. * Support pre-sales discussions, solution proposals, estimations, and customer presentations. ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [From event streaming to event sourcing 101](https://www.wearedevelopers.com/videos/91-from-event-streaming-to-event-sourcing-101) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Introducing JSON Structure](https://www.wearedevelopers.com/videos/100219-introducing-json-structure) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)