> Markdown version of [/jobs/ext/2608224-big-data-pyspark-lead-engineer-vice-president](https://www.wearedevelopers.com/jobs/ext/2608224-big-data-pyspark-lead-engineer-vice-president). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Big Data PySpark Lead Engineer - Vice President - **Company:** Citi - **Location:** Jersey City, NJ, United States - **Experience:** Expert - **Salary:** $142,320.0 - $213,480.0 - **Contract:** Permanent contract - **Skills:** Data Analysis, Software Applications, CA Workload Automation Ae, Big Data, Cloud Computing, Cloudera Impala, Computer Programming, Data Architecture, Information Engineering, Data Infrastructure, Data Warehousing, Software Debugging, Distributed Data Store, Distributed Systems, Apache Hadoop, Hadoop Distributed File System, Apache Hive, Systems Analysis, Job Scheduling, Scala (Programming Language), Shell Script, Software Engineering, SQL Databases, Sqoop, Data Streaming, Unstructured Data, Freeform SQL, Data Ingestion, System Availability, Apache Spark, Software Application Programming, Data Lakes, Pyspark, Apache Kafka, Data Delivery, Stream Processing, Data Pipelines - **Published:** August 2, 2026 - **Apply:** https://www.dice.com/job-detail/ca8100c8-dbbe-42e0-8604-93d42e2057b2 ## About the Role * Hands-on expertise in PySpark and Big Data processing, with the ability to build and optimize distributed data workflows at scale. * Practical knowledge of the Hadoop ecosystem, including Hive, HDFS, Sqoop, Spark, Impala, and Scala, applied in a production environment. * Proficiency in complex SQL query development for data analysis, transformation, and validation across large datasets. * Solid understanding of distributed systems architecture and how data flows across interconnected processing layers. * Demonstrated knowledge of data modelling and data design, with familiarity in data warehouse concepts and dimensional modelling techniques. * Competence in shell scripting and job scheduling using Autosys or equivalent workflow automation tools. * Strong analytical and problem-solving ability, with a track record of working independently to diagnose and resolve complex data engineering challenges. * Clear and effective communication skills, with the ability to articulate technical concepts to both technical and non-technical audiences. Beneficial Skills & Qualifications * Familiarity with streaming data platforms such as Apache Kafka or equivalent real-time data processing technologies. * Exposure to cloud-based Big Data environments and modern data lake architectures. * Experience working in financial services or regulated industries where data quality and governance are critical., * 6 -10 years of relevant experience in Apps Development or systems analysis role * Extensive experience system analysis and in programming of software applications * Experience in managing and implementing successful projects * Subject Matter Expert (SME) in at least one area of Applications Development * Ability to adjust priorities quickly as circumstances dictate * Demonstrated leadership and project management skills * Consistently demonstrates clear and concise written and verbal communication Education: * Bachelor?s degree/University degree or equivalent experience * Master?s degree preferred ## Description The Big Data PySpark Lead Engineer is a senior level position responsible for establishing and implementing new or revised application systems and programs in coordination with the Technology team. The overall objective of this role is to lead applications systems analysis and programming activities. * Partner with multiple management teams to ensure appropriate integration of functions to meet goals as well as identify and define necessary system enhancements to deploy new products and process improvements * Resolve variety of high impact problems/projects through in-depth evaluation of complex business processes, system processes, and industry standards * Provide expertise in area and advanced knowledge of applications programming and ensure application design adheres to the overall architecture blueprint * Utilize advanced knowledge of system flow and develop standards for coding, testing, debugging, and implementation * Develop comprehensive knowledge of how areas of business, such as architecture and infrastructure, integrate to accomplish business goals * Provide in-depth analysis with interpretive thinking to define issues and develop innovative solutions * Serve as advisor or coach to mid-level developers and analysts, allocating work as necessary * Appropriately assess risk when business decisions are made, demonstrating particular consideration for the firm\'s reputation and safeguarding Citigroup, its clients and assets, by driving compliance with applicable laws, rules and regulations, adhering to Policy, applying sound ethical judgment regarding personal behavior, conduct and business practices, and escalating, managing and reporting control issues with transparency., * Build and maintain scalable data pipelines using PySpark within a Big Data environment to process and transform large volumes of structured and unstructured data. * Design and develop solutions across the Hadoop ecosystem ? including Hive, HDFS, Sqoop, Spark, Impala, and Scala ? to enable efficient data ingestion, processing, and storage. * Develop and manage real-time and batch data workflows using streaming data platforms, ensuring high availability and low-latency data delivery. * Write complex SQL queries to extract, validate, and analyze data across distributed systems, supporting data-driven decision-making. * Design and implement data models and data architecture patterns aligned with data warehouse principles, ensuring scalability, accuracy, and consistency. * Automate pipeline scheduling and orchestration using shell scripting and Autosys, reducing manual intervention and improving operational reliability. * Independently identify, assess, and resolve technical risks and data issues in a timely manner, maintaining system integrity across the data platform. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Data Analyst Salary in the UK](https://www.wearedevelopers.com/magazine/278-data-analyst-salary-in-the-uk)