Pyspark/Python Data Engineer

Tata Consultancy Services Limited
Irving, TX, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$100,000.0 - $120,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Airflow Amazon Web Services Amazon S3 Continuous Integration Information Engineering Extract Transform Load (ETL) Data Transformation Data Warehousing Software Debugging Distributed Computing Environment Distributed Systems
+11 more
Apache Hive Python (Programming Language) Performance Tuning Standard Sql Data Processing Snowflake Git Pyspark Information Technology Software Version Control Data Pipelines

Requirements

Do you have experience in Technical architecture?, Do you have a Bachelor’s degree?, Must Have Technical/Functional Skills We are looking for a skilled PySpark Data Engineer with strong hands-on experience in PySpark and Python to design, build, and optimize scalable data processing pipelines. The ideal candidate will have practical experience working with distributed data processing and a solid foundation in writing efficient, production-grade Python code Required Technical Skills

  • Strong hands-on experience in PySpark (Spark SQL, DataFrame API)

  • Advanced proficiency in Python (data processing, performance tuning, modular coding)

  • Solid understanding of ETL design patterns and data pipeline architecture

  • Good working knowledge of SQL for data transformation and analysis

  • Experience with data processing in distributed environments

Preferred Skills (Good to Have)

  • Experience with cloud platforms (AWS preferred - S3, Glue, EMR or equivalent services)

  • Familiarity with workflow orchestration tools such as Airflow or similar schedulers

  • Exposure to data warehousing concepts (e.g., Snowflake or similar platforms)

  • Knowledge of code versioning (Git) and CI/CD practices

Experience

  • 3-8 years of experience in Data Engineering / PySpark development

  • Proven hands-on project experience in PySpark + Python

Roles & Responsibilities

  • Design, develop, and maintain ETL/ELT pipelines using PySpark

  • Write optimized and scalable PySpark transformations using DataFrames and Spark SQL

  • Develop reusable and efficient Python-based data processing components

  • Ensure data quality, integrity, and performance across pipelines

  • Perform debugging, performance tuning, and optimization of PySpark jobs, Qualifications : BACHELOR OF COMPUTER SCIENCE

Benefits & conditions

(part of Tata group) 3.93.9 out of 5 stars Irving, TX $100,000 - $120,000 a year

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · WWC Europe 2026

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all