Data Engineer - Onsite

MSYS Inc.
Wilmington, DE, United States
about 1 month ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary to permanent
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Amazon Web Services Amazon S3 Big Data Cloud Computing Cloud Engineering Information Engineering Data Infrastructure Data Transformation Distributed Computing Environment Distributed Data Store Distributed Systems
+20 more
Python (Programming Language) Microsoft Message Queuing Cloud Services Data Processing Data Ingestion Snowflake Apache Spark Infrastructure as Code (IaC) Event Driven Architecture Pyspark AWS Aurora Apache Kafka Data Management Terraform Stream Processing Data Pipelines AWS EKS Amazon Elastic Mapreduce (EMR) Service Stack Databricks

Job description

We are seeking a highly skilled Data Engineer to join a team responsible for building and enhancing a large-scale decision platform that drives customer-focused business decisions. The ideal candidate will have strong expertise in Spark, Java, AWS, and large-scale data processing environments., * Design, develop, and optimize scalable data pipelines for ingesting and processing high-volume datasets.

  • Build and maintain distributed data processing solutions using Apache Spark and Java.
  • Process and manage 20M+ to 40M+ daily data records efficiently.
  • Develop batch and real-time data processing workflows.
  • Work with event-driven architectures and streaming platforms.
  • Collaborate with cross-functional teams to enhance data platform capabilities.
  • Implement and manage cloud-native solutions within AWS environments.
  • Utilize Infrastructure as Code (IaC) methodologies using Terraform.

Requirements

  • 5 to 10 years of experience.
  • Strong experience with Apache Spark (Must Have)
  • Strong experience with Java (Must Have)
  • Hands-on experience with AWS Cloud Services (Must Have)
  • Experience with Terraform for Infrastructure as Code (Must Have)
  • Experience building data ingestion and data transformation pipelines
  • Experience working with large-scale distributed data environments
  • Knowledge of real-time and batch data processing architectures

Preferred skills:

  • Python
  • Apache Kafka

Cloud & Technology Stack:

  • Apache Spark
  • Java
  • AWS EKS
  • AWS EMR
  • AWS S3
  • AWS Aurora
  • AWS MSK
  • AWS SNS
  • AWS SQS
  • Apache Kafka
  • Terraform, * Strong background in data engineering and distributed computing.
  • Experience handling high-volume data platforms.
  • Comfortable working in cloud-native, event-driven architectures.
  • Excellent problem-solving and analytical skills.

Must have:

  • Pyspark
  • Snowflake
  • Databricks
  • AWS

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:49 min

Container hosting options available on Amazon Web Services

Federico Fregosi · World Congress 2022

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

Videos

See all

Related articles

See all