> Markdown version of [/jobs/ext/1976332-java-developer-with-spark](https://www.wearedevelopers.com/jobs/ext/1976332-java-developer-with-spark). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Java developer with Spark - **Company:** Propertyvalue Quantum Technologies Llc - **Location:** Berkeley Heights, NJ, United States - **Experience:** Expert - **Salary:** $180,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Big Data, Continuous Integration, Data Architecture, Data Governance, Extract Transform Load (ETL), Database Queries, Distributed Systems, Memory Management, Fault Tolerance, Performance Tuning, Workflow Management Systems, Parquet, Apache Yarn, Apache Spark, Containerization, Data Lakes, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Apache Flink, Avro, Apache Kafka, Data Management, Stream Processing, Data Pipelines - **Published:** August 7, 2026 - **Apply:** https://www.dice.com/job-detail/86b2ac18-cffa-4fde-b5f3-82f5f4be7ab7 ## About the Role Bachelor s or master s degree in computer science, Engineering, or related field 7+ years of professional Java development experience 5+ years hands-on experience with Apache Spark in production environments Expert-level understanding of distributed systems: fault tolerance, data locality, shuffle mechanics, resource management Proven track record designing systems processing terabyte+ scale data Strong SQL skills and deep familiarity with columnar storage formats (Parquet, ORC, Avro, Delta Lake/Iceberg) Experience with cluster managers (YARN, Kubernetes) and cloud-managed Spark Proficiency with Kafka Strong grasp of CI/CD, containerization, and infrastructure-as-code practices Preferred Qualifications Experience with Flink or other stream-processing frameworks Familiarity with data governance, lineage, and quality frameworks Experience with workflow orchestration at scale Background in system design for multi-tenant or multi-region data platforms Prior experience leading a team or acting as a technical lead Soft Skills / Leadership Excellent communication able to explain technical tradeoffs to non-technical stakeholders Strong mentorship and coaching ability Comfortable driving ambiguous, cross-team technical initiatives ## Description Architect and build scalable, fault-tolerant data pipelines using Apache Spark (Java) Lead design of batch and streaming ETL/ELT systems handling large data volumes Deep-dive performance tuning: partitioning strategy, memory management, shuffle/skew optimization, job cost reduction Set coding standards and lead code/design reviews across the team Drive technical decisions on data architecture, storage formats, and pipeline orchestration Mentor mid-level and junior engineers; act as a technical escalation point Partner with product, analytics, and platform teams to translate requirements into scalable systems Own production reliability on-call ownership, incident response, root-cause analysis for pipeline failures Evaluate and introduce new tools/frameworks where they improve the system Contribute to capacity planning and cost optimization for cluster infrastructure ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Parquet, Delta, Iceberg & Ducklake - An introduction for developers](https://www.wearedevelopers.com/videos/100075-parquet-delta-iceberg-ducklake-an-introduction-for-developers) - [Flex your Energy: Building a Cloud-Native Platform for Renewable Energy Communities](https://www.wearedevelopers.com/videos/1990-flex-your-energy-building-a-cloud-native-platform-for-renewable-energy-communities) - [From event streaming to event sourcing 101](https://www.wearedevelopers.com/videos/91-from-event-streaming-to-event-sourcing-101) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Green Cloud Computing](https://www.wearedevelopers.com/videos/592-green-cloud-computing) ## Related Articles - [Top 10 Java Libraries](https://www.wearedevelopers.com/magazine/364-top-10-java-libraries) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [93 Java Interview Questions You Should Prepare For](https://www.wearedevelopers.com/magazine/14-93-java-interview-questions-you-should-prepare-for) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers)