> Markdown version of [/jobs/ext/527668-biba-practice-cloud-data-lead](https://www.wearedevelopers.com/jobs/ext/527668-biba-practice-cloud-data-lead). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # BIBA Practice - Cloud Data Lead - **Company:** Hexaware Technologies - **Location:** United States - **Experience:** Expert - **Salary:** $151,840.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Application Programming Interfaces (APIs), Airflow, Amazon S3, Automation of Tests, Catalyst (Software), Cloud Computing, Cloud Database, Cloud Engineering, Cloud Storage, Profiling, Concurrent Computing, Information Engineering, Data Files, Data Infrastructure, Extract Transform Load (ETL), Memory Management, Elasticsearch, Github, Apache Hadoop, Hadoop Distributed File System, Monitoring of Systems, Apache Hive, Java Virtual Machine (JVM), Python (Programming Language), Message Broker, NoSQL, Performance Tuning, Prometheus, Software Engineering, Data Streaming, Systems Integration, Parquet, Data Logging, Data Processing, Data Storage Technologies, Feature Engineering, Data Ingestion, Apache Yarn, Grafana, Apache Spark, Containerization, Pyspark, Gitlab-ci, Kubernetes, Information Technology, Low Latency, Avro, Apache Kafka, Apache Nifi, Data Pipelines, Jenkins - **Published:** June 5, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=2bd4b7b9df019bdf ## About the Role Do you have experience in Yarn (JavaScript package manager)?, We are seeking a Senior Spark Engineer with strong Java expertise to design, develop, and operate high-performance, production-scale data processing pipelines. The role focuses on Apache Spark-based batch and streaming solutions, robust ETL, performance tuning, and close collaboration with data engineering, data science, and platform teams. Technical Skills (Must-Have): Experience with Scala or Python (PySpark) for cross-language integrations. Java (primary): language proficiency, performance profiling, GC tuning. Apache Spark: job design, RDD/DataFrame/Dataset APIs, Catalyst optimizer understanding. Structured Streaming: exactly-once semantics, watermarking, state management. Data storage: Hive, Parquet/ORC, Avro, schema evolution best practices. Messaging & ingestion: Apache Kafka (producers/consumers), Connectors. Orchestration & CI/CD: Airflow, Jenkins/GitHub Actions/GitLab CI or equivalent. Containerization/cluster deployment: Yarn, Kubernetes experience for Spark on K8s. Monitoring & observability: Prometheus/Grafana, ELK/EFK stack or Cloud-native equivalents. 5+ years of software engineering experience with at least 3+ years building production systems using Apache Spark. Strong Java development skills (Java 8+); solid understanding of concurrent programming, memory management, and JVM tuning. Production experience with Spark Core, Spark SQL, and Structured Streaming. Hands-on experience with the Hadoop ecosystem components (HDFS, YARN, Hive) or cloud object storage (S3/GCS/Azure Blob). Experience integrating with Kafka or other message brokers for real-time ingestion., Bachelor's or Master's degree in Computer Science, Engineering, or equivalent practical experience. ## Description We kindly request that you refrain from posting any of Hexaware's job openings on LinkedIn, as doing so may be perceived as competition. Additionally, we ask that there be no use of hashtags, mentions of Hexaware, or references to its customers on any online platforms. Your cooperation in maintaining this confidentiality and professionalism is greatly appreciated., Own design and development of scalable data pipelines using Apache Spark for batch and streaming workloads. Implement Spark applications in Java (primary) and integrate with the broader data platform (HDFS/S3, Hive, Kafka, relational and NoSQL stores). Optimize Spark jobs for performance, memory usage, and resource efficiency; troubleshoot production issues and reduce job failures/latency. Develop reusable libraries, frameworks, and abstractions to accelerate data engineering work. Implement data ingestion, transformation, and enrichment patterns, ensuring data quality, schema evolution handling, and idempotence. Integrate Spark workloads with orchestration and scheduling systems (Airflow/Elasticsearch/Nifi or equivalent). Build and maintain CI/CD pipelines, automated tests (unit/integration), and deployment practices for data applications. Collaborate with data scientists to productionize models and feature engineering pipelines. Drive observability and monitoring for Spark jobs (metrics, logging, alerting). Mentor and review work of mid/junior engineers; participate in architecture and design reviews. Ensure security, governance, and compliance requirements are met for data processing. ## Related Videos - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [From event streaming to event sourcing 101](https://www.wearedevelopers.com/videos/91-from-event-streaming-to-event-sourcing-101) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [NoSQL Data Modeling for Front-end Developers](https://www.wearedevelopers.com/videos/297-nosql-data-modeling-for-front-end-developers) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Top 10 Java Libraries](https://www.wearedevelopers.com/magazine/364-top-10-java-libraries) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [93 Java Interview Questions You Should Prepare For](https://www.wearedevelopers.com/magazine/14-93-java-interview-questions-you-should-prepare-for)