PySpark Developer

Integrated Technology Corporation
United States
10 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Agile Methodology Airflow Amazon Web Services Apache HTTP Server Big Data Continuous Integration Data Architecture Information Engineering Data Infrastructure Extract Transform Load (ETL) Data Transformation
+18 more
DevOps Identity and Access Management Python (Programming Language) Metadata Performance Tuning Query Optimization Standard Sql Cloudera SQL Databases Unstructured Data Apache Spark Git Amazon Relational Database Service Data Lakes Pyspark Information Technology Apache Kafka Cloudwatch

Job description

Seeking a Senior PySpark Developer with strong experience in PySpark, Cloudera CDP, Apache Iceberg, Airflow, SQL, and AWS to support a large-scale cloud transformation and data modernization initiative.

Requirements

  • 10+ years IT experience; 6+ years in Data Engineering/Big Data
  • Strong hands-on PySpark/Spark development
  • Experience with Cloudera Data Platform (CDP)
  • Strong knowledge of Apache Iceberg and Data Lake/Lakehouse architecture
  • Experience with Apache Airflow
  • Strong SQL, ETL/ELT, and data modeling skills
  • Large-scale structured/unstructured data processing
  • Git, CI/CD, and DevOps practices
  • Performance tuning, partitioning, and query optimization

Preferred:

  • AWS: S3, Glue, Lambda, EMR, ECS/EKS, RDS, IAM, CloudWatch
  • Kafka/event-driven processing
  • Python/Java development
  • Data quality, metadata, and governance
  • Healthcare/Medicaid/Claims/Insurance experience
  • Agile and cloud migration/modernization experience
  • AWS or Cloudera certification

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:41 min

Visualizing the complex developer journey for JVM ecosystems

Bobur Umurzokov · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all