BIBA Practice - Cloud Data Lead

Hexaware Technologies
United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$151,840.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Application Programming Interfaces (APIs) Airflow Amazon S3 Automation of Tests Catalyst (Software) Cloud Computing Cloud Database Cloud Engineering Cloud Storage Profiling Concurrent Computing
+40 more
Information Engineering Data Files Data Infrastructure Extract Transform Load (ETL) Memory Management Elasticsearch Github Apache Hadoop Hadoop Distributed File System Monitoring of Systems Apache Hive Java Virtual Machine (JVM) Python (Programming Language) Message Broker NoSQL Performance Tuning Prometheus Software Engineering Data Streaming Systems Integration Parquet Data Logging Data Processing Data Storage Technologies Feature Engineering Data Ingestion Apache Yarn Grafana Apache Spark Containerization Pyspark Gitlab-ci Kubernetes Information Technology Low Latency Avro Apache Kafka Apache Nifi Data Pipelines Jenkins

Job description

We kindly request that you refrain from posting any of Hexaware’s job openings on LinkedIn, as doing so may be perceived as competition. Additionally, we ask that there be no use of hashtags, mentions of Hexaware, or references to its customers on any online platforms. Your cooperation in maintaining this confidentiality and professionalism is greatly appreciated., Own design and development of scalable data pipelines using Apache Spark for batch and streaming workloads. Implement Spark applications in Java (primary) and integrate with the broader data platform (HDFS/S3, Hive, Kafka, relational and NoSQL stores). Optimize Spark jobs for performance, memory usage, and resource efficiency; troubleshoot production issues and reduce job failures/latency. Develop reusable libraries, frameworks, and abstractions to accelerate data engineering work. Implement data ingestion, transformation, and enrichment patterns, ensuring data quality, schema evolution handling, and idempotence. Integrate Spark workloads with orchestration and scheduling systems (Airflow/Elasticsearch/Nifi or equivalent). Build and maintain CI/CD pipelines, automated tests (unit/integration), and deployment practices for data applications. Collaborate with data scientists to productionize models and feature engineering pipelines. Drive observability and monitoring for Spark jobs (metrics, logging, alerting). Mentor and review work of mid/junior engineers; participate in architecture and design reviews. Ensure security, governance, and compliance requirements are met for data processing.

Requirements

Do you have experience in Yarn (JavaScript package manager)?, We are seeking a Senior Spark Engineer with strong Java expertise to design, develop, and operate high-performance, production-scale data processing pipelines. The role focuses on Apache Spark-based batch and streaming solutions, robust ETL, performance tuning, and close collaboration with data engineering, data science, and platform teams. Technical Skills (Must-Have): Experience with Scala or Python (PySpark) for cross-language integrations. Java (primary): language proficiency, performance profiling, GC tuning. Apache Spark: job design, RDD/DataFrame/Dataset APIs, Catalyst optimizer understanding. Structured Streaming: exactly-once semantics, watermarking, state management. Data storage: Hive, Parquet/ORC, Avro, schema evolution best practices. Messaging & ingestion: Apache Kafka (producers/consumers), Connectors. Orchestration & CI/CD: Airflow, Jenkins/GitHub Actions/GitLab CI or equivalent. Containerization/cluster deployment: Yarn, Kubernetes experience for Spark on K8s. Monitoring & observability: Prometheus/Grafana, ELK/EFK stack or Cloud-native equivalents.

5+ years of software engineering experience with at least 3+ years building production systems using Apache Spark. Strong Java development skills (Java 8+); solid understanding of concurrent programming, memory management, and JVM tuning. Production experience with Spark Core, Spark SQL, and Structured Streaming. Hands-on experience with the Hadoop ecosystem components (HDFS, YARN, Hive) or cloud object storage (S3/GCS/Azure Blob). Experience integrating with Kafka or other message brokers for real-time ingestion., Bachelor’s or Master’s degree in Computer Science, Engineering, or equivalent practical experience.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:09 min

Evaluating mature stream processing frameworks for production systems

Soroosh Khodami Soroosh Khodami · WWC 2024

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

3:02 min

Audience Q&A on data formats and engine tradeoffs

Matthias Niehoff Matthias Niehoff · WWC Europe 2026

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

3:16 min

Terminology differences between relational and NoSQL databases

Tim Faulkes · LIVE

Videos

See all

Related articles

See all