Hadoop Developer

IBA InfoTech Inc.
Newark, NJ, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours

Tech stack

Data Analysis Big Data Data Integration Data Mining Relational Databases Apache Hadoop Hadoop Distributed File System Apache HBase Apache Hive JSON Python (Programming Language) Node.Js
+10 more
Object-Oriented Software Development Scala (Programming Language) SQL Databases Sqoop Data Streaming Pyspark Apache Flume Information Technology Apache Kafka Data Pipelines

Requirements

  • Need a strong Scala/PySpark programmer. Python experience also.
  • Expert SQL background
  • 8+ years of experience in data analysis, data modelling and implementation of enterprise class systems spanning Big Data, Data Integration, Object Oriented programming and Advanced Analytics
  • Excellent understanding of Hadoop architecture and different demons of Hadoop clusters which include Job Tracker, Task Tracker, Name Node and Data Node
  • Good understanding of Data Mining and techniques
  • Experience in importing and exporting data from RDBMS to HDFS, Hive tables and HBase by using Sqoop
  • Experience in importing streaming data into HDFS using flume sources, and flume sinks and transforming the data using flume interceptors
  • Exposure on usage of Apache Kafka develop data pipeline of logs as a stream of messages using producers and consumers
  • Knowledge of HBase and Json.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on ibainfotech.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

1:28 min

Building shared Java modules and analyst targeting platforms

Chris Heilmann +2 · LIVE

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters · WWC 2025

1:09 min

Evaluating mature stream processing frameworks for production systems

Soroosh Khodami Soroosh Khodami · WWC 2024

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all