Senior Data Engineer

EXL SERVICE
United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Compensation
$130,000.0 - $150,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Amazon S3 Apache HTTP Server Basic Access Authentication Big Data Information Engineering Data Files Extract Transform Load (ETL) Relational Databases Database Queries File Systems Amazon DynamoDB
+26 more
Hadoop Distributed File System JSON Python (Programming Language) MongoDB NoSQL OAuth Oracle (Applications) Test Data Extensible Markup Language (XML) Parquet Flask (Web Framework) Apache Spark Boto3 AWS Lambda Fastapi Pandas Pyspark Information Technology Avro AWS Glue Apache Kafka Cloudwatch Restful APIs GPT Data Pipelines Docker

Job description

  • Build and test data processing applications using PySpark and Python
  • Develop data pipelines using AWS Glue ETL or EMR
  • Create AWS Lambda functions using Python (pandas, json, requests, awswrangler)
  • Work with data from:
  • Relational databases (Oracle)
  • NoSQL databases (DynamoDB, MongoDB)
  • File systems (S3, HDFS)
  • Implement event-driven pipelines using Kafka
  • Develop REST APIs using FastAPI or Flask
  • Implement basic authentication (OAuth2 / JWT) for APIs
  • Build workflow orchestration pipelines using Step Functions and EventBridge
  • Work with big data file formats such as Parquet, Avro, ORC, JSON
  • Optimize Spark jobs using standard techniques (partitioning, joins, etc.)
  • Use Glue Crawlers to catalog datasets
  • Monitor and troubleshoot jobs using CloudWatch
  • Support deployment using Docker containers

Requirements

Do you have experience in XML?, Do you have a Bachelor’s degree?, * Hands-on experience with Python and PySpark

  • Experience with AWS services:
  • S3, Glue, Lambda, EMR, Step Functions, EventBridge, Athena
  • Experience with Kafka integration
  • Strong SQL skills (writing complex queries)
  • Experience working with data file formats (Parquet, Avro, ORC, JSON, XML)
  • Experience using Python libraries (pandas, requests, boto3)
  • Experience building REST APIs (FastAPI or Flask)

Experience

  • 4+ years of experience in Data Engineering or related field
  • Bachelor’s degree in Computer Science or related field (or equivalent)

Base Compensation Range: $130,000 - $150,000

The posted range is the hiring range for this role - a subset of the broader range available to employees over time - and reflects base salary across our national hiring scale. Final offers are based on several factors, including the candidate’s skills and experience, internal pay equity, work location, market conditions for the role, and the specific scope and responsibilities of the position. The top of the range is reserved for candidates who notably exceed the requirements; the lower end applies to those with less experience or fewer preferred qualifications. For positions based in higher-cost zones (e.g., California, New York, New Jersey), actual compensation may exceed the posted range; your recruiter will share specifics during the process.

Responsibilities: The client is specifically looking for candidates with strong hands-on experience in the following technologies:

  • Building data pipelines using Python and PySpark on AWS Glue, EMR, and Lambda
  • Developing and securing RESTful APIs (FastAPI) deployed on Docker/EKS, with OAuth2/JWT-based authentication
  • Hands-on experience with Apache Iceberg tables for CDC and latest snapshot handling
  • Designing event-based pipelines using Apache Kafka / MSK for data consumption and publishing
  • Ability to lead and communicate complex technical designs, and leverage Copilot/GPT for agentic coding across the stack

Qualifications: The client is specifically looking for candidates with strong hands-on experience in the following technologies:

  • Building data pipelines using Python and PySpark on AWS Glue, EMR, and Lambda
  • Developing and securing RESTful APIs (FastAPI) deployed on Docker/EKS, with OAuth2/JWT-based authentication
  • Hands-on experience with Apache Iceberg tables for CDC and latest snapshot handling
  • Designing event-based pipelines using Apache Kafka / MSK for data consumption and publishing
  • Ability to lead and communicate complex technical designs, and leverage Copilot/GPT for agentic coding across the stack

Benefits & conditions

3.73.7 out of 5 stars United States Hybrid work $130,000 - $150,000 a year - Full-time

About the company

We are looking for a Data Engineer with strong skills in Python and PySpark to design and build data solutions for a Fortune 500 client. The role focuses on building data pipelines and integrations on AWS as part of an enterprise data lake platform.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

3:02 min

Audience Q&A on data formats and engine tradeoffs

Matthias Niehoff Matthias Niehoff · WWC Europe 2026

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · WWC 2024

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters · WWC 2025

1:52 min

Customizing block storage tiers and formats

Ricardo Sueiras Sueiras · LIVE

Videos

See all

Related articles

See all