> Markdown version of [/jobs/ext/1893075-java-spark-engineer](https://www.wearedevelopers.com/jobs/ext/1893075-java-spark-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Java Spark Engineer - **Company:** Aventine software - **Location:** Berkeley Heights, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Big Data, Data Architecture, Extract Transform Load (ETL), Database Queries, Distributed Systems, Memory Management, Fault Tolerance, Performance Tuning, Parquet, Apache Yarn, Apache Spark, Data Lakes, Kubernetes, Information Technology, Avro, Apache Kafka, Data Pipelines - **Published:** August 2, 2026 - **Apply:** https://www.dice.com/job-detail/f8467eb8-89c0-4412-9a5e-493597e6bec4 ## About the Role * Bachelor's or Master's degree in Computer Science, Engineering, or related field * 7+ years of professional Java development experience * 5+ years hands-on experience with Apache Spark in production environments * Expert-level understanding of distributed systems: fault tolerance, data locality, shuffle mechanics, resource management * Proven track record designing systems processing terabyte+ scale data * Strong SQL skills and deep familiarity with columnar storage formats (Parquet, ORC, Avro, Delta Lake/Iceberg) * Experience with cluster managers (YARN, Kubernetes) and cloud-managed Spark * Proficiency with Kafka ## Description * Architect and build scalable, fault-tolerant data pipelines using Apache Spark (Java) * Lead design of batch and streaming ETL/ELT systems handling large data volumes * Deep-dive performance tuning: partitioning strategy, memory management, shuffle/skew optimization, job cost reduction * Set coding standards and lead code/design reviews across the team * Drive technical decisions on data architecture, storage formats, and pipeline orchestration * Mentor mid-level and junior engineers; act as a technical escalation point * Partner with product, analytics, and platform teams to translate requirements into scalable systems * Own production reliability - on-call ownership, incident response, root-cause analysis for pipeline failures * Evaluate and introduce new tools/frameworks where they improve the system * Contribute to capacity planning and cost optimization for cluster infrastructure ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Parquet, Delta, Iceberg & Ducklake - An introduction for developers](https://www.wearedevelopers.com/videos/100075-parquet-delta-iceberg-ducklake-an-introduction-for-developers) - [From event streaming to event sourcing 101](https://www.wearedevelopers.com/videos/91-from-event-streaming-to-event-sourcing-101) - [Green Cloud Computing](https://www.wearedevelopers.com/videos/592-green-cloud-computing) - [Let's Get Aggregated: Custom UDAFs in Spark ](https://www.wearedevelopers.com/videos/1649-let-s-get-aggregated-custom-udafs-in-spark) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top 10 Java Libraries](https://www.wearedevelopers.com/magazine/364-top-10-java-libraries) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)