> Markdown version of [/jobs/ext/1323722-lead-data-engineer](https://www.wearedevelopers.com/jobs/ext/1323722-lead-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead Data Engineer - **Company:** UFS LLC - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon S3, Apache HTTP Server, Information Systems, Databases, Information Engineering, Data Governance, Document-Oriented Databases, Python (Programming Language), PostgreSQL, Machine Learning, Microsoft SQL Server, SQL Databases, File Transfer Protocol (FTP), Data Ingestion, Debezium, Information Technology, Apache Kafka, Presto, Data Pipelines - **Published:** July 17, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=ec49f4a8c516eedd ## About the Role To perform this job successfully, an individual must be able to perform each essential duty satisfactorily. The requirements listed below are representative of the knowledge, skill, and/or ability required. * 8-12+ years in data engineering with end-to-end ownership of ingestion through serving, and 2+ years in a lead or senior role * Strong Python and expert SQL; rigorous data modeling for analytics * Hands-on lakehouse experience (Iceberg/Delta/Hudi or equivalent) and modern transformation tooling * Built reliable pipelines from messy operational and transactional source systems * Comfort with CDC mechanics and the realities of pulling from databases you do not control Core Technologies * Languages: Python, SQL (deep) * Lakehouse & catalog: Apache Iceberg; Polaris / Nessie / Lakekeeper * Transform & query: dbt; Trino / Presto / DuckDB * CDC & streaming: Debezium (SQL Server CDC, Postgres logical replication), Kafka / Redpanda * Orchestration: Dagster (or Airflow) * Storage: S3 / MinIO * SQL Server and PostgreSQL data modeling, pgvector (or equivalent) Nice to Have * Experience with financial or core-banking data, or FFIEC / Call Report data specifically * Strong SQL Server familiarity * Data contracts, lineage, and governance practices Education and/or Experience * Bachelor's degree in computer science, mathematics, information systems, or a related field, or equivalent hands-on experience * Experience in the financial services industry or a regulated data environment strongly preferred Work Structure & Expectations * Full-time role combining ongoing pipeline operations with initiative-based lakehouse build-out and new bank onboarding * Close collaboration with AI/ML, platform engineering, and security teams; on-call rotation covering data pipeline reliability Physical Demands The physical demands described here are representative of those that must be met by an employee to successfully perform the essential functions of this job. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions. While performing the duties of this job, the employee is regularly required to sit and use hands to finger, handle, or touch objects, tools, or controls. The employee frequently is required to talk or hear. The employee is occasionally required to stand; walk; and stoop, kneel, crouch, or crawl. The employee must occasionally lift and/or move up to 10 pounds, usually waist high, up to 50 feet away. Specific vision abilities required by this job include close vision and the ability to adjust focus. ## Description * Design the lakehouse: Apache Iceberg (or similar technology) on object storage, a catalog for table management and per-bank isolation, dbt models, and a query engine * Build secure, least-privilege ingestion from bank systems - log-based CDC where permitted, with query-based and batch/SFTP fallbacks, plus an in-bank collector pattern * Own data modeling for the semantic and metric layer (deposits, concentration, uninsured exposure, asset quality, and peer groups) * Handle schema drift, data quality, and reconciliation; make ingestion observable and recoverable * Partner with the AI/ML team on the structured-query path and with Security on PII classification at landing, in alignment with regulatory data-handling requirements * Document data lineage, transformation logic, and access controls to support audit and exam readiness * Define and enforce data contracts, quality thresholds, and alerting for pipeline failures Core Competencies * End-to-end ownership of ingestion-through-serving pipelines, with a bias toward reliability and observability * Rigorous data modeling for analytics - semantic layers, metric definitions, and reconcilable outputs * Security and compliance mindset: PII handling, least-privilege access, and data governance aligned to regulatory guidance * Cross-functional partnership with AI/ML and platform engineering to deliver governed, queryable data products Key Performance Indicators (KPIs) * Data freshness and pipeline reliability - SLAs met for data ingestion and bank-core feeds * Data quality score across key metrics versus source reconciliation * Time to onboard a new bank's data environment, from kickoff to queryable lakehouse * PII classification coverage at landing and zero unauthorized data-access incidents * Semantic layer adoption - percentage of assistant queries resolved via governed metrics versus ad hoc SQL ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Python-Based Data Streaming Pipelines Within Minutes](https://www.wearedevelopers.com/videos/1233-python-based-data-streaming-pipelines-within-minutes) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Swapping a Data Warehouse at Runtime: Zero-Downtime Migration Without Changing a Single Client](https://www.wearedevelopers.com/videos/100311-swapping-a-data-warehouse-at-runtime-zero-downtime-migration-without-changing-a-single-client) ## Related Articles - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)