ETL Developer

TCS - Tata Consultancy Services
Plano, TX, United States
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Airflow Big Data Continuous Integration Data Governance Extract Transform Load (ETL) Data Systems Apache Hadoop Hadoop Distributed File System MapReduce Apache Hive Apache Oozie Performance Tuning
+16 more
Query Optimization Standard Sql Unstructured Data Cloud Platform System Apache Yarn Git Data Lakes Pyspark Real Time Data Apache Kafka Bitbucket Data Management Tez (Software) Data Pipelines Jenkins Databricks

Job description

Experteer Overview In this role you will design and optimize large-scale data platforms for global markets. You will build real-time and batch data pipelines, focusing on PySpark, Kafka, and Hadoop ecosystems, while advancing Databricks Lakehouse architectures. You will work closely with quants, risk teams, and product owners to deliver governed, high-performance data solutions for regulatory, trading, and analytics workloads. The opportunity centers on shaping scalable, compliant data platforms that support mission-critical financial functions. You will mentor engineers and contribute to strong engineering practices. Compensation / Benefits * Design, develop, and optimize PySpark ETL pipelines on on-prem Hadoop clusters and cloud environments * Build high-volume ingestion frameworks using Kafka for real-time data * Tuning and managing Hadoop components: HDFS, YARN, MapReduce/Tez, Oozie/Airflow * Develop high-performance Hive data models for regulatory reporting and risk processing * Architect and implement Bronze/Silver/Gold layer modeling in Databricks Lakehouse * Apply Delta Lake best practices (file management, Z-Ordering, CDF, schema evolution, ACID) * Create reusable ingestion, cleansing, transformation, and consumption frameworks across Lakehouse layers * Enable governance, lineage, and auditability using cataloging tools (Unity Catalog or equivalent) * Collaborate with quants, product owners, risk tech, and business users; participate in agile ceremonies * Mentor junior engineers and promote strong engineering practices across teams Tasks * 10-13 years of hands-on Big Data engineering experience * Expert in PySpark (optimizations, partitioning, broadcasting) * Expert in Kafka (producer/consumer design, schema registry, streaming ETLs) * Strong Hadoop ecosystem knowledge (HDFS, YARN, MapReduce/Tez, Oozie/Airflow) * Advanced Hive skills (query tuning, TEZ, partitioning) * Extensive Databricks Lakehouse experience (Bronze/Silver/Gold, Delta Lake optimizations) * Experience with data quality frameworks on Lakehouse and handling structured/unstructured data * Experience in Global Markets, Risk, Treasury, Trade Surveillance, or Regulatory Reporting * Strong SQL on TB/PB-scale datasets * Experience with CI/CD practices (Git, Jenkins, Bitbucket) * Familiarity with governance/catalog tools for lineage and auditability * Experience with on-prem and cloud big data platforms Key requirements * Discretionary annual incentive * Medical coverage * Parental leaves * Commuter benefits * Certification & training reimbursement * Vacation & holidays

Requirements

Jenkins, and implement Bronze/Silver/Gold layer modeling in Databricks Lakehouse * Apply Delta Lake best practices (file management, Z-Ordering, CDF, schema evolution, ACID) * Create reusable ingestion, cleansing, transformation, and consumption frameworks across Lakehouse layers * Enable governance, lineage, and auditability using cataloging tools (Unity Catalog or equivalent) * Collaborate with quants, product owners, risk tech, and business users; participate in agile ceremonies * Mentor junior engineers and promote strong engineering practices across teams Tasks * 10-13 years of hands-on Big Data engineering experience * Expert in PySpark (optimizations, partitioning, broadcasting) * Expert in Kafka (producer/consumer design, schema registry, streaming ETLs) * Strong Hadoop ecosystem knowledge (HDFS, YARN, MapReduce/Tez, Oozie/Airflow) * Advanced Hive skills (query tuning, TEZ, partitioning) * Extensive Databricks Lakehouse experience (Bronze/Silver/Gold, Delta Lake optimizations) * Experience with data quality frameworks on Lakehouse and handling structured/unstructured data * Experience in Global Markets, Risk, Treasury, Trade Surveillance, or Regulatory Reporting * Strong SQL on TB/PB-scale datasets * Experience with CI/CD practices (Git, Jenkins, Bitbucket) * Familiarity with governance/catalog tools for lineage and auditability * Experience with on-prem and cloud big data platforms Key requirements * Discretionary annual incentive * Medical coverage * Parental leaves * Commuter benefits * Certification & training reimbursement * Vacation & holidays

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

47 sec

Building modern data pipelines for legacy exports

Dr. Alexander Wachtel Dr. Alexander Wachtel +1 · WWC 2025

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier · WWC 2023

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

57 sec

Extracting API schemas automatically during continuous integration builds

Axel Barbier · WWC 2023

Videos

See all

Related articles

See all