Data Engineer

Siri InfoSolutions Inc
West Chester, PA, United States
7 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Databases Data Validation Extract Transform Load (ETL) Relational Databases Distributed Computing Environment Python (Programming Language) Performance Tuning SQL Databases Teradata SQL Snowflake Apache Spark
+5 more
Git Pyspark Star Schema Software Version Control Data Pipelines

Job description

Pipeline Development: Design and implement robust ETL/ELT pipelines using PySpark and SQL to ingest data from diverse sources including APIs, flat files, and relational databases Data Modeling: Develop and manage complex data models (e.g., Star/Snowflake schemas) and maintain Fact and Dimension tables within Snowflake Performance Optimization: Monitor and tune Snowflake queries and Spark jobs to optimize performance, reduce latency, and manage computational costs Data Quality & Integrity: Implement automated data validation frameworks and testing procedures to ensure a “single source of truth” and high data reliability Operations: Troubleshoot production pipeline issues, manage version control via Git

Requirements

Must Have Technical/Functional Skills Snowflake: Deep expertise in Snowflake PySpark: Strong hands-on experience using Apache Spark with Python for distributed data processing and transformation. Database & SQL Knowledge: Advanced proficiency in Teradata

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg · LIVE

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

56 sec

Introduction to analytical data formats for software developers

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

Videos

See all

Related articles

See all