Python/Big Data Developer

IBA InfoTech Inc.
Charlotte, NC, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Data Analysis Big Data Information Engineering Apache Hadoop Apache Hive Python (Programming Language) Simple Data Format SQL Databases Parquet Data Processing Data Ingestion
+4 more
Apache Spark Data Lakes Pyspark Information Technology

Job description

  • In-depth understanding and knowledge of Hadoop and Spark architecture and RDD transformation

Requirements

  • Proven experience in developing solutions using Spark architecture and PySpark for data engineering pipelines, transformation, and aggregation of data from a variety of sources into the data lake.
  • At least 3 or more years of relevant experience in developing PySpark programs using APIs. Expertise in different file formats like parquet, ORC.
  • Experience with troubleshooting, fine-tuning Spark and python based applications for scalability and performance.
  • Experience in designing hive tables to handle velocity, variety and to handle huge volumes.
  • Experience in data ingestion, processing and analyzing data using Spark/SQL from disparate sources.
  • Knowledge in using Spark-Submit and Spark UI. Experience in creating and then performing operations on Spark RDD.
  • Experience in creating Spark Data Frames from RDD, HIVE and Parquet files and then performing Joins and Aggregations on Dataframes.
  • Experience in processing data from Python and other API modules.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on ibainfotech.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:50 min

How Parquet metadata enables efficient data reading

Matthias Niehoff Matthias Niehoff · WWC Europe 2026

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all