PySpark Developer

Sage IT Inc
Irving, TX, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours
Job source

Tech stack

Big Data Cloud Computing Computer Programming Extract Transform Load (ETL) Data Transformation Data Security Data Warehousing Relational Databases DevOps Distributed Systems Apache Hadoop Apache Hive
+10 more
Python (Programming Language) Standard Sql SAS (Software) SQL Databases Data Streaming Delivery Pipeline Apache Spark Pyspark Apache Kafka Data Pipelines

Requirements

Experience with big data processing and distributed computing systems like Spark.

Implement ETL pipelines and data transformation processes.

Ensure data quality and integrity in all data processing workflows.

Troubleshoot and resolve issues related to PySpark applications and workflows.

Understand source, dependencies and data flow from converted PySpark code.

Strong programming skills in Python and SQL.

Experience with big data technologies like Hadoop, Hive, and Kafka.

Understanding of data warehousing concepts and relational databases like SQL.

Demonstrate and document code lineage.

Integrate PySpark code with frameworks such as Ingestion Framework, DataLens, etc.,

Ensure compliance with data security, privacy regulations, and organizational standards.

Knowledge of CI/CD pipelines and DevOps practices.

Strong problem-solving and analytical skills.

Excellent communication and leadership abilities.

Qualifications:

6+ years of experience in big data development, Hadoop , Hive & Spark framework.

Good to have experience in SAS.

Strong Python, PySpark Development and SQL knowledge.

Certification in big data or cloud technologies is preferred.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

1:41 min

Visualizing the complex developer journey for JVM ecosystems

Bobur Umurzokov · LIVE

Videos

See all

Related articles

See all