Senior Data Engineer _ Python with Spark

Tata Consultancy Services Limited
Jersey City, NJ, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$100,000.0 - $115,000.0
Working hours
Regular working hours
Job source

Tech stack

Adobe InDesign Airflow Amazon Web Services Microsoft Azure Big Data Cloud Computing Data as a Services Data Architecture Data Validation Information Engineering Extract Transform Load (ETL) Data Security
+23 more
Data Systems Data Warehousing DevOps Distributed Systems Apache Hadoop Apache Hive JSON Python (Programming Language) Performance Tuning Standard Sql SQL Databases Unstructured Data Parquet Data Processing Cloud Platform System Data Ingestion Apache Spark Pyspark Information Technology Avro Integration Frameworks Non-relational Database Data Pipelines

Job description

  • Design, develop, and maintain scalable data pipelines using Python and Spark.

  • Build and optimize ETL/ELT workflows for processing large volumes of structured and unstructured data.

  • Work closely with data analysts, data scientists, and business stakeholders to understand data requirements.

  • Develop efficient and reusable data processing frameworks and components.

  • Perform data validation, cleansing, and transformation to ensure quality and consistency.

  • Optimize and tune Spark jobs and SQL queries for performance and scalability.

  • Collaborate with DevOps and platform teams to deploy and manage data solutions in cloud environments.

  • Ensure data security, governance, and compliance standards are met.

  • Troubleshoot production issues, perform root cause analysis, and implement fixes.

  • Participate in design discussions, contribute to data architecture decisions, and promote best practices.

  • Work in an Agile environment, supporting sprint activities and continuous improvement initiatives.

Requirements

Do you have experience in Spark implementation?, Do you have a Bachelor’s degree?, Must Have Technical/Functional Skills

  • Strong hands-on experience in Python programming for data engineering and data processing.

  • Extensive experience with Apache Spark (PySpark) for large-scale data processing and distributed computing.

  • Strong knowledge of SQL and experience working with relational and non-relational databases.

  • Experience in building and maintaining ETL/ELT pipelines for data ingestion and transformation.

  • Good understanding of data warehousing concepts, data modeling, and data architecture.

  • Experience working with big data technologies such as Hadoop ecosystem, Hive, or similar platforms.

  • Familiarity with cloud platforms (AWS, Azure, or GCP) and related data services.

  • Hands-on experience with data pipeline orchestration tools such as Airflow or similar.

  • Knowledge of data formats such as Parquet, Avro, JSON, and CSV.

  • Experience with performance tuning and optimization of Spark jobs and data pipelines.

  • Strong problem-solving skills and ability to work with cross-functional teams., Qualifications : BACHELOR OF COMPUTER SCIENCE

Benefits & conditions

3.93.9 out of 5 stars Jersey City, NJ $100,000 - $115,000 a year, Pulled from the full job description

  • Pet insurance
  • Health insurance
  • Vision insurance
  • Dental insurance
  • Commuter assistance, Discretionary Annual Incentive. Comprehensive Medical Coverage: Medical & Health, Dental & Vision, Disability Planning & Insurance, Pet Insurance Plans. Family Support: Maternal & Parental Leaves. Insurance Options: Auto & Home Insurance, Identity Theft Protection. Convenience & Professional Growth: Commuter Benefits & Certification & Training Reimbursement. Time Off: Vacation, Time Off, Sick Leave & Holidays. Legal & Financial Assistance: Legal Assistance, 401K Plan, Performance Bonus, College Fund, Student Loan Refinancing. Salary Range: $100,000 - $115,000 a year

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

3:02 min

Audience Q&A on data formats and engine tradeoffs

Matthias Niehoff Matthias Niehoff · WWC Europe 2026

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters · WWC 2025

Videos

See all

Related articles

See all