Sr Hadoop+Spark(scala) Data Engineer

VIRTUES INCORPORATED
Irving, TX, United States
12 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$100,696.0 - $110,269.0
Working hours
Regular working hours
Job source

Tech stack

Batch Processing Big Data BigQuery Cloud Computing Computer Programming System Configuration Data Architecture Data Validation Data Integration Data Transformation Data Stores Data Systems
+40 more
Data Warehousing Software Debugging Distributed File Systems Distributed Computing Environment Distributed Data Store Fault Tolerance Apache Hadoop Hadoop Distributed File System MapReduce Apache HBase Apache Hive MapR (Big Data) Apache Oozie Performance Tuning Raw Data Release Management Cloudera Scala (Programming Language) Sqoop Data Streaming Backup and Restore Apache Zookeeper Data Processing Google Cloud Data Ingestion Apache Spark Apache Pig Data Layers Event Driven Architecture Data Lakes Apache Flume Information Technology Enterprise Integration Apache Kafka Spark Streaming Data Management Data Delivery Stream Processing Software Version Control Data Pipelines

Job description

We are seeking an experienced Senior Big Data Engineer with strong expertise in the Hadoop ecosystem, Apache Spark, Scala, Kafka, and large-scale data processing. The ideal candidate will have extensive hands-on experience designing, developing, and supporting scalable batch and real-time data pipelines across multiple data platforms.

This role requires strong technical capabilities in data ingestion, distributed data processing, data transformation, data modeling, and cloud-based big data technologies. Experience with Hadoop platform administration & Google Cloud Platform (GCP) Bigquery will be a plus.

Key Responsibilities

Big Data Engineering & Data Pipeline Development

Design, develop, implement, and maintain scalable, high-performance data ingestion and processing pipelines using Hadoop ecosystem technologies.

Develop and manage data pipelines supporting multiple data ingestion and processing modes, including:

  • Batch data processing

  • Real-time/event-driven data processing

  • Streaming data processing

Build robust data ingestion solutions that:

  • Read data from multiple structured, semi-structured, and unstructured data sources.

  • Ingest batch data and real-time event streams, including Kafka events.

  • Perform complex data validation, cleansing, enrichment, and transformation.

  • Deliver processed data to target data stores, curated data layers, publishing zones, and downstream endpoints.

Develop and optimize Apache Spark applications using Scala for large-scale distributed data processing.

Design and implement Kafka-centric event processing and real-time data pipelines.

Develop streaming data transformation logic using Apache Spark Streaming and/or Spark Structured Streaming.

Build and maintain scalable batch processing solutions using Apache Spark.

Develop data processing and analytical solutions using HiveQL, Pig Latin, HBase, and custom MapReduce programs.

Develop data transformation and integration processes to move data from raw data zones to curated and published data warehouse layers.

Collaborate with data architects, application teams, business stakeholders, and platform teams to translate functional requirements into scalable technical solutions.

Hadoop Ecosystem & Platform Engineering

Work extensively with Hadoop ecosystem technologies, including:

  • HDFS

  • MapReduce

  • Hive

  • Pig

  • Sqoop

  • HBase

  • ZooKeeper

  • Oozie

  • Apache Spark

  • Scala

  • Flume/Flume NG

  • Kafka

  • Hue

Apply strong knowledge of Hadoop architecture and core components, including:

  • NameNode

  • DataNode

  • HDFS architecture

  • JobTracker

  • TaskTracker

  • MapReduce programming and execution paradigm

Install, configure, integrate, and support Hadoop ecosystem components within Cloudera-based environments.

Work with distributed storage and processing frameworks to ensure scalability, reliability, fault tolerance, and high performance.

Monitor and optimize data pipeline performance, resource utilization, throughput, and processing efficiency.

Data Modeling & Data Integration

Design and implement data models, data transformation processes, and detailed technical designs.

Develop scalable data integration solutions to move data across raw, staging, curated, and publishing layers.

Support data warehouse and data lake architectures and ensure high-quality, reliable, and timely data delivery.

Implement data quality, validation, reconciliation, and error-handling mechanisms within data pipelines.

Ensure data solutions are scalable, maintainable, reusable, and aligned with enterprise data architecture standards.

Requirements

Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering, Data Science, or a related technical discipline.

7+ years of strong hands-on experience in Hadoop framework and the broader Hadoop ecosystem.

6+ years of hands-on experience developing data ingestion and integration solutions across multiple data platforms.

5+ years of strong hands-on experience in Apache Spark with Scala-based distributed data processing.

5+ years of experience in data modeling, data transformation, detailed technical design, and data integration.

Strong experience designing and developing large-scale batch and real-time data pipelines.

Strong experience with HiveQL, Pig Latin, HBase, and custom MapReduce programming.

Experience developing and managing Kafka-centric event-driven data pipelines.

Strong understanding of batch processing, stream processing, and event-driven architecture.

Hands-on experience with Spark Streaming and/or Spark Structured Streaming.

Experience installing and configuring Cloudera Hadoop ecosystem components, including Hive, HBase, ZooKeeper, Oozie, Spark, Sqoop, Flume, Pig, and Hue.

Strong understanding of Hadoop architecture, HDFS, distributed storage, and MapReduce concepts.

Strong analytical, problem-solving, debugging, and performance-tuning skills.

Excellent communication and collaboration skills.

Preferred / Nice-to-Have Skills:

Hadoop Platform Administration, GCP Bigquery

The following Hadoop administration and platform engineering skills are highly desirable:

Experience providing end-to-end Hadoop administration and production support.

Experience with Hadoop infrastructure setup, software installation, configuration, upgrades, patching, monitoring, troubleshooting, and maintenance.

Experience administering Hadoop distributions and platforms such as:

  • Cloudera

  • MapR

  • Hortonworks

Experience installing, configuring, and managing Hadoop ecosystem components, including Hive, Pig, HBase, ZooKeeper, Oozie, Spark, and related services.

Experience managing and monitoring HDFS, distributed file systems, and Hadoop clusters.

Experience managing, monitoring, scheduling, and troubleshooting MapReduce and distributed processing jobs.

Experience with cluster capacity planning, resource management, health monitoring, and operational support.

Experience automating operational activities using scripting, including:

  • Backup and restore processes

  • Cluster monitoring

  • Health checks

  • Maintenance activities

  • Operational reporting

Experience with version control, change management, release management, incident management, problem management, and root-cause analysis.

Key Competencies

Strong expertise in distributed data processing and big data architecture.

Deep understanding of batch, real-time, streaming, and event-driven data processing.

Strong hands-on programming skills in Scala and distributed data engineering frameworks.

Ability to design scalable, fault-tolerant, and high-performance data solutions.

Strong technical troubleshooting and root-cause analysis capabilities.

Ability to work independently while collaborating effectively with cross-functional teams.

Strong ownership, attention to detail, and commitment to data quality and operational excellence.

Benefits & conditions

$100,696.26 - $110,268.61 a year - Full-time, Contract

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:27 min

Explaining query execution overhead and caching limitations in BigQuery

Adnan Rahic · JS Congress

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · WWC Europe 2026

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all