PySpark/Spark Data Pipeline

Computer Enterprises, Inc.
Pittsburgh, PA, United States
3 months ago
Apply on indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Business Logic Data Validation Dataspaces Software Debugging Linux Apache Hadoop Python (Programming Language) Object-Oriented Software Development Performance Tuning SQL Databases Apache Spark
+5 more
Git Pandas Pyspark Data Pipelines Jenkins

Job description

Participate on daily scrum calls for upcoming change release implementations - assigned work, status, questions, collaboration. Coding, testing, advancements, and modernization as a software developer. Key Responsibilities

  • Participate on daily scrum calls for upcoming change release implementations - assigned work, status, questions, collaboration.
  • Coding, testing, advancements, and modernization as a software developer

Requirements

Do you have experience in System performance optimization?, * Strong Spark + SQL transformation skills

  • Rigorous data validation & reconciliation mindset
  • Ability to translate STM/business logic into working pipelines
  • Experience operating in structured, governed data ecosystems
  • Practical debugging and performance optimization skills
  • Python, SQL, Git and Linux skills are required.
  • experience in working with wrangling data and building data pipelines

Preferred Skills

  • Preferrable experience in Jenkins, Apache Spark & Hadoop (Pandas is also good too), and object oriented programming.
  • Flex Skills Knowledge of AI space, acceleration enhancements and improvements.
  • Experience with Co-Pilot or AI code enablement functionality.

INDGEN

#ZR

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

1:34 min

Bringing diverse skills to industrial data science roles

Katja Träumner

Videos

See all

Related articles

See all