Java Spark Engineer

Propertyvalue Quantum Technologies Llc
Westlake, United States of America
2 days ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
$ 125K

Job location

Remote
Westlake, United States of America

Tech stack

Java
Big Data
Continuous Integration
Data Architecture
Data Governance
ETL
Database Queries
Distributed Systems
Memory Management
Fault Tolerance
Performance Tuning
Workflow Management Systems
Parquet
Apache Yarn
Spark
Containerization
Data Lake
Kubernetes
Infrastructure Automation Frameworks
Information Technology
Apache Flink
Avro
Kafka
Data Management
Stream Processing
Data Pipelines

Job description

Architect and build scalable, fault-tolerant data pipelines using Apache Spark (Java)

Lead design of batch and streaming ETL/ELT systems handling large data volumes

Deep-dive performance tuning: partitioning strategy, memory management, shuffle/skew optimization, job cost reduction

Set coding standards and lead code/design reviews across the team

Drive technical decisions on data architecture, storage formats, and pipeline orchestration

Mentor mid-level and junior engineers; act as a technical escalation point

Partner with product, analytics, and platform teams to translate requirements into scalable systems

Own production reliability on-call ownership, incident response, root-cause analysis for pipeline failures

Evaluate and introduce new tools/frameworks where they improve the system

Contribute to capacity planning and cost optimization for cluster infrastructure

Requirements

Bachelor s or Master s degree in Computer Science, Engineering, or related field

7+ years of professional Java development experience

5+ years hands-on experience with Apache Spark in production environments

Expert-level understanding of distributed systems: fault tolerance, data locality, shuffle mechanics, resource management

Proven track record designing systems processing terabyte+ scale data

Strong SQL skills and deep familiarity with columnar storage formats (Parquet, ORC, Avro, Delta Lake/Iceberg)

Experience with cluster managers (YARN, Kubernetes) and cloud-managed Spark

Proficiency with Kafka

Strong grasp of CI/CD, containerization, and infrastructure-as-code practices

Preferred Qualifications

Experience with Flink or other stream-processing frameworks

Familiarity with data governance, lineage, and quality frameworks

Experience with workflow orchestration at scale

Background in system design for multi-tenant or multi-region data platforms

Prior experience leading a team or acting as a technical lead

Soft Skills / Leadership

Excellent communication able to explain technical tradeoffs to non-technical stakeholders

Strong mentorship and coaching ability

Comfortable driving ambiguous, cross-team technical initiatives

Apply for this position