Data Engineer - AWS & PySpark

Cliff Services Inc
Dallas, TX, United States
12 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Amazon S3 Big Data Data as a Services Extract Transform Load (ETL) Data Transformation Data Systems Distributed Computing Environment Python (Programming Language) Cloud Services Standard Sql Data Processing
+3 more
Cloud Platform System Pyspark Data Pipelines

Job description

  • Design, develop, and maintain scalable data pipelines using PySpark and AWS.
  • Develop and optimize large-scale data processing and ETL workflows.
  • Work with AWS cloud services to build reliable and high-performance data solutions.
  • Perform data transformation, cleansing, validation, and integration.
  • Optimize PySpark jobs for performance, scalability, and cost efficiency.
  • Collaborate with data architects, analysts, developers, and business stakeholders.
  • Troubleshoot data pipeline issues and ensure data quality and reliability.
  • Follow engineering best practices for code development, testing, deployment, and documentation.

Requirements

We are looking for an experienced Data Engineer with strong hands-on expertise in AWS and PySpark. The ideal candidate must have prior professional experience working with Banking domain and be capable of developing scalable data pipelines and data processing solutions in a cloud environment., * Strong hands-on experience with PySpark.

  • Strong experience with AWS cloud services.
  • Experience developing ETL/data pipelines and processing large datasets.
  • Strong Python and SQL skills.
  • Experience with data transformation, integration, and data quality.
  • Mandatory: Prior Capital One project/client experience.
  • Strong communication and problem-solving skills.

Preferred

  • Experience with AWS data services such as S3, Glue, EMR, Lambda, Redshift, or similar.
  • Experience with distributed data processing and cloud-based data platforms.
  • Financial services/banking domain experience.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann +3 · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

Videos

See all

Related articles

See all