> Markdown version of [/jobs/ext/1970735-hadoop-developer](https://www.wearedevelopers.com/jobs/ext/1970735-hadoop-developer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Hadoop Developer - **Company:** EdgeAll Inc - **Location:** Jersey City, NJ, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Airflow, Architectural Patterns, Cloudera Impala, Information Engineering, Data Governance, Extract Transform Load (ETL), Apache Hadoop, Hadoop Distributed File System, Apache Hive, Python (Programming Language), Apache Oozie, Performance Tuning, Systems Development Life Cycle, Scala (Programming Language), SQL Databases, Workflow Management Systems, Apache Yarn, Apache Spark, Technical Debt, Data Lakes, Pyspark, Information Technology, Apache Kafka, Spark Streaming, Data Management, Stream Processing, Control M - **Published:** August 7, 2026 - **Apply:** https://www.dice.com/job-detail/2b46ce91-bd20-4370-b128-ea7406d59005 ## About the Role Primary Skill: Data Engineering, Platform Engineering or architecture roles. Deep Expertise in Pyspark. Experience: 10+ yrs Roles & Responsibilities Bachelor s or master s degree in computer science or related field. Required Hard and Soft Skills / Experience Deep Expertise in PySpark, including performance tuning and optimization Strong python development experience in large-scale distributed environment Solid knowledge of Hadoop ecosystem (HDFS,Hive/Impala, YARN) Proven experience designing and governing enterprise, regulatory facing data platforms. Expertise in designing data lakes, ELT/ETL pipelines, batch and real time data processing solution Proficiency in programming languages such as Java, Scala and SQL Strong understanding of non-functional requirements and production support models Clear written and verbal communication skills with ability to influence across organizations Preferred Skills / Experience Financial services experience, particularly in Market Surveillance, AML, Fraud, Or Risk Technology Experience supporting regulatory or audit facing platforms Kafka and Spark Structured streaming exposure Familiarity with Orchestration tools(Airflow,Control-M,Oozie) Knowledge of data governance, lineage, and data quality controls ## Description Define target state architecture and strategic roadmap for surveillance data processing platforms leveraging Hadoop,Spark, Pyspark and Python Establish standard architectural patterns for alert generation, enrichment, aggregation and reporting pipelines. Ensure architecture aligns with enterprise technology standards and Surveillance control expectation Solution Design and Delivery Support Translate surveillance business requirements(e.g, market misconduct detection, regulatory coverage) into scalable technical designs Review and approve detailed technical designs, ensuring alignment with functional intent, regulatory requirements, and architectural standards. Provide hands-on architectural guidance to engineering teams during development, testing and implementation Design and standardization of Spark/Pyspark frameworks supporting surveillance alert generation and enrichment Modernization and Optimization of large-scale Hadoop surveillance workloads to improve performance, stability and control coverage Implementation of enterprise-consistent architecture patterns enabling audit readiness, lineage, and regulatory traceability Support end-to-end delivery across the SDLC, minimizing rework and technical debt Architecture Governance & Leadership Participate in and lead architecture review forums and design walkthroughs Mentor senior engineers and promote adoption of Standard framework and best practices Influence enterprise surveillance platform strategy through architecture governance Documentation & Communication Maintain high-quality documentation, including business requirement, data definitions and process flows Proactively communicate risks, dependencies and potential timeline impact to stakeholders ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Brewing Tea over the Internet](https://www.wearedevelopers.com/videos/697-brewing-tea-over-the-internet) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top 10 Java Libraries](https://www.wearedevelopers.com/magazine/364-top-10-java-libraries) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk)