Senior Specialist - Data Engineering

LTM Inc
Charlotte, NC, United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Amazon Web Services Microsoft Azure Big Data Information Engineering Data Governance Distributed Computing Environment Apache Hadoop Hadoop Distributed File System Apache Hive Performance Tuning Sqoop Unstructured Data
+5 more
Data Processing Data Ingestion Apache Spark Pyspark Data Pipelines

Job description

Experteer Overview As a Data Engineer, you will design, develop and maintain scalable data pipelines in on-premise Hadoop environments. You will work with PySpark, HDFS and Hive to process large-scale data and optimize Spark jobs. You’ll troubleshoot data pipelines, enforce data quality practices, and collaborate with cross-functional teams to enable reliable data-driven decisions. This role focuses on hands-on implementation and performance tuning in a fast-paced setting, with exposure to cloud platforms as a plus. Compensation / Benefits * design and maintain scalable data pipelines * develop and optimize Spark jobs (PySpark) * work with HDFS and Hive for data processing and warehousing * data ingestion and integration using Sqoop * troubleshoot data pipeline issues and performance bottlenecks * enforce data quality governance and best practices * support onprem Hadoop ecosystem architectures and tuning * collaborate with stakeholders in fast-paced environments Tasks * 5+ years of experience with HDFS, Hive and Spark * strong hands-on experience in the Hadoop ecosystem (on-premise) * expertise in PySpark for large-scale data processing * experience building and optimizing Spark jobs * hands-on data ingestion experience with Sqoop * good understanding of distributed data processing and big data concepts * ability to design, develop and maintain scalable data pipelines * experience with large volumes of structured and unstructured data * strong problem-solving skills and independence in fast-paced environments * exposure to cloud platforms (AWS/Azure/GCP) is a plus Key requirements *

Requirements

_ of experience with HDFS, Hive and Spark * strong hands-on experience in the Hadoop ecosystem (on-premise) * expertise in PySpark for large-scale data processing * experience building and optimizing Spark jobs * hands-on data ingestion experience with Sqoop * good understanding of distributed data processing and big data concepts * ability to design, develop and maintain scalable data pipelines * experience with large volumes of structured and unstructured data * strong problem-solving skills and independence in fast-paced environments * exposure to cloud platforms (AWS/Azure/GCP) is a plus Key requirements *

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

Videos

See all

Related articles

See all