EMR/Spark SQL & Job Query-Tuning Engineer

Amazon.com, Inc.
Seattle, WA, United States
17 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$91,000.0 - $152,000.0
Working hours
Regular working hours

Tech stack

Big Data Catalyst (Software) Cloud Computing Directed Acyclic Graph (Directed Graphs) Information Engineering Serialization Distributed Computing Environment Distributed Systems Apache Hive Performance Tuning Query Optimization SQL Databases
+8 more
Parquet Data Processing Apache Spark Pyspark Information Technology Integration Frameworks AWS Data Analytics Amazon Elastic Mapreduce (EMR)

Job description

STAFFXPERT LLC is seeking an EMR/Spark SQL & Job Query-Tuning Engineer on behalf of our client in Seattle, WA. This role is ideal for a Spark performance expert who specializes in optimizing PySpark, Spark SQL, and Hive workloads within large-scale data environments. The successful candidate will focus on improving application-level performance by tuning queries, optimizing execution plans, reducing memory consumption, minimizing shuffle operations, and enhancing overall job efficiency on Amazon EMR. Key Responsibilities

  • Optimize and tune PySpark, Spark SQL, and Hive SQL jobs to improve performance, scalability, and resource utilization.
  • Analyze Spark execution plans and DAGs to identify and resolve performance bottlenecks.
  • Design and implement efficient partitioning, caching, and data processing strategies.
  • Reduce shuffle overhead, spill events, executor memory pressure, and job execution times.
  • Optimize join strategies, including broadcast joins, sort-merge joins, and other Spark execution techniques.
  • Troubleshoot data skew, serialization issues, and distributed processing inefficiencies.
  • Collaborate with data engineering and platform teams to improve workload performance and reliability.
  • Monitor, benchmark, and continuously enhance large-scale data processing jobs.
  • Recommend and implement best practices for Spark and EMR performance optimization., Hi, Role: EMR/Spark SQL & Job Query-Tuning Engineer Location:Redmond,WA skills Required: Application/code-level optimization of PySpark/Spark/Hive SQL jobs; query tuning, D…
  • 20 hours ago
  • Apply easily +

Requirements

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent professional experience.
  • Strong hands-on experience with PySpark, Spark SQL, Hive SQL, and Amazon EMR.
  • Proven expertise in Spark performance tuning, query optimization, and execution plan analysis.
  • Deep understanding of Spark internals, distributed computing concepts, and data processing frameworks.
  • Experience optimizing large-scale data workloads and addressing memory, partitioning, and shuffle-related challenges.
  • Proficiency in analyzing Spark UI metrics and troubleshooting performance issues.
  • Strong problem-solving skills and ability to work in a fast-paced environment.

Preferred Qualifications

  • Knowledge of Adaptive Query Execution (AQE), Catalyst Optimizer, and Tungsten Engine.
  • Experience working with Parquet, ORC, and other columnar data formats.
  • Familiarity with AWS data services and cloud-based big data platforms.
  • Experience supporting enterprise-scale analytics or data engineering environments.

About the company

STAFFXPERT LLC is a leading talent solutions and technology staffing firm connecting skilled professionals with innovative organizations across diverse industries. We are committed to delivering exceptional opportunities and helping top talent advance their careers through impactful projects and long-term partnerships. Job Summary

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:34 min

Capabilities of the Apache Spark processing engine

Ayon Roy · LIVE

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

2:50 min

How Parquet metadata enables efficient data reading

Matthias Niehoff Matthias Niehoff · World Congress 2026 Europe

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all