Pyspark Developer

TSR, Inc
Tampa, FL, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Third Normal Form BigQuery Databases Directed Acyclic Graph (Directed Graphs) Data Architecture Extract Transform Load (ETL) Data Structures Data Systems Data Vault Modeling Data Warehousing Relational Databases Software Design Patterns
+14 more
Distributed Systems Memory Management Apache Hadoop Apache Hive Python (Programming Language) PostgreSQL Oracle (Applications) SQL Databases Snowflake Apache Spark Pyspark Optimization Algorithms Star Schema Data Pipelines

Requirements

  • PySpark Mastery: Production-level expertise in Apache Spark using Python (PySpark)
  • Must understand Spark internals (DAGs, shuffling, memory management, and optimization techniques)
  • Data Modeling: Proven track record of building complex data models from scratch (Star/Snowflake schemas, Data Vault, or 3NF)
  • Database & SQL: Expert-level proficiency in SQL
  • Extensive hands-on experience with massive relational databases (e.g., Oracle, PostgreSQL) and modern data warehouses/lakes (e.g., Snowflake, BigQuery, or Hive/Hadoop)
  • Systems Design: Clear understanding of distributed systems processing, ETL/ELT design patterns, and enterprise data warehousing principles
  • Communication: Demonstrated ability to translate complex technical concepts into clear, concise language for non-technical stakeholders and business leaders
  • We are seeking a highly experienced PySpark Developer to lead the design, architecture, and development of mission-critical data pipelines and enterprise data models
  • As a senior technical leader, you will bridge the gap between complex business requirements and highly scalable data architecture
  • The ideal candidate possesses deep expertise in PySpark, advanced SQL optimization, and enterprise data modeling
  • You will not only be a hands-on technical contributor but also serve as an architectural guide, mentoring junior developers, establishing best practices, and ensuring that data solutions are highly performant, resilient, and aligned with Client
  • Lead the transition from legacy data structures to modern, scalable cloud/hybrid

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg · LIVE

2:30 min

Leveraging BigQuery ML for scalable SQL-based segmentation experiments

Julian Joseph · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

4:01 min

Managing application isolation via pluggable database models

Wei Hu Wei Hu · WWC 2022

3:27 min

Explaining query execution overhead and caching limitations in BigQuery

Adnan Rahic · JS Congress

Videos

See all

Related articles

See all