Data Engineer (Pyspark)

Cliff Services Inc
Westwood, MA, United States
3 months ago
Apply on dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Amazon S3 Extract Transform Load (ETL) SQL Databases Parquet Snowflake Apache Spark Data Lakes Pyspark Apache Kafka Amazon Elastic Mapreduce (EMR)

Requirements

Key Skills: Apache Spark (PySpark/Scala), SQL, Redshift/Snowflake, ETL/ELT, Data Lake (S3, Parquet, Iceberg), Kafka, CDC pipelines, AWS EMR

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann +3 · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:50 min

How Parquet metadata enables efficient data reading

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

1:41 min

Visualizing the complex developer journey for JVM ecosystems

Bobur Umurzokov · LIVE

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all