Senior Data Engineer - Identity Resolution & Large-Scale Data Engineering

Amazon.com, Inc.
Bellevue, WA, United States
18 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Microsoft Azure Big Data Customer Data Management Data Deduplication Information Engineering Data Systems Distributed Computing Environment Python (Programming Language) Machine Learning Operational Databases Performance Tuning
+17 more
Standard Sql Web Applications Azure Data Factory Large Language Models Apache Spark Generative AI Scalability Testing Pyspark Low Latency Apache Kafka Cosmos DB Spark Streaming Virtual Agents Data Pipelines Serverless Computing User Identification Databricks

Job description

We are seeking an experienced Senior Data Engineer to design, build, and optimize production-grade data solutions supporting Identity Resolution (IDR) and large-scale customer data initiatives., Design and optimize “large-scale data processing solutions capable of operating at 1B+ record scale”. Develop scalable data pipelines using Azure Databricks, PySpark, Python, and ADF”. Optimize Spark workloads including large-scale joins, partitioning, shuffling, data skew, memory, and compute utilization. Design and improve solutions for Identity Resolution, Entity Resolution, Record Linkage, Deduplication, or similar high-volume matching problems. Evaluate technical approaches against production data volumes and infrastructure constraints. Identify and communicate scalability and environment constraints early and recommend viable alternatives. Perform performance and scalability testing using production-like workloads. Build reliable data pipelines with appropriate validation, error handling, monitoring, and recovery mechanisms. Collaborate with architects, engineers, data scientists, and stakeholders to deliver solutions from design through production.

Requirements

The ideal candidate will have proven experience delivering data solutions at billion-record scale, with strong expertise in Azure Databricks, PySpark, Python, Azure Data Factory, and distributed data processing. The role requires strong production-scale engineering judgment-the ability to identify technical and scalability constraints early, evaluate alternatives quickly, and deliver tested, scalable, and production-ready solutions. Experience with Identity Resolution or comparable high-throughput matching systems is highly preferred., 7+ years of Data Engineering experience. Proven experience working with very large-scale datasets, preferably 1B+ records. Strong hands-on experience with Azure Databricks and PySpark. Strong Python and SQL skills. Experience with Azure Data Factory (ADF) and Azure data services. Strong understanding of distributed processing and Spark performance optimization. Experience delivering production-grade, scalable data solutions. Ability to identify technical constraints early and make timely technical decisions. Strong problem-solving and communication skills. Preferred Qualifications Experience with Identity Resolution / Entity Resolution / Record Linkage / Deduplication / Fuzzy Matching. Experience with Cosmos DB and low-latency APIs. Experience with Azure Functions or Azure Web Apps. Experience with Event Hub, Kafka, or Spark Structured Streaming. ML experience related to matching or customer data. Exposure to Generative AI, LLMs, RAG, or Agentic AI. Telecommunications or large-scale customer data experience.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

6:35 min

Evaluating blob storage against cosmos db for unstructured data

Menaka Baskerpillai · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all