Java Spark Engineer

Aventine software
Berkeley Heights, United States
9 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Big Data Data Architecture Extract Transform Load (ETL) Database Queries Distributed Systems Memory Management Fault Tolerance Performance Tuning Parquet Apache Yarn Apache Spark
+6 more
Data Lakes Kubernetes Information Technology Avro Apache Kafka Data Pipelines

Job description

  • Architect and build scalable, fault-tolerant data pipelines using Apache Spark (Java)

  • Lead design of batch and streaming ETL/ELT systems handling large data volumes

  • Deep-dive performance tuning: partitioning strategy, memory management, shuffle/skew optimization, job cost reduction

  • Set coding standards and lead code/design reviews across the team

  • Drive technical decisions on data architecture, storage formats, and pipeline orchestration

  • Mentor mid-level and junior engineers; act as a technical escalation point

  • Partner with product, analytics, and platform teams to translate requirements into scalable systems

  • Own production reliability - on-call ownership, incident response, root-cause analysis for pipeline failures

  • Evaluate and introduce new tools/frameworks where they improve the system

  • Contribute to capacity planning and cost optimization for cluster infrastructure

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related field

  • 7+ years of professional Java development experience

  • 5+ years hands-on experience with Apache Spark in production environments

  • Expert-level understanding of distributed systems: fault tolerance, data locality, shuffle mechanics, resource management

  • Proven track record designing systems processing terabyte+ scale data

  • Strong SQL skills and deep familiarity with columnar storage formats (Parquet, ORC, Avro, Delta Lake/Iceberg)

  • Experience with cluster managers (YARN, Kubernetes) and cloud-managed Spark

  • Proficiency with Kafka

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:02 min

Audience Q&A on data formats and engine tradeoffs

Matthias Niehoff Matthias Niehoff · WWC Europe 2026

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

2:50 min

How Parquet metadata enables efficient data reading

Matthias Niehoff Matthias Niehoff · WWC Europe 2026

1:52 min

Customizing block storage tiers and formats

Ricardo Sueiras Sueiras · LIVE

1:41 min

Visualizing the complex developer journey for JVM ecosystems

Bobur Umurzokov · LIVE

Videos

See all

Related articles

See all