> Markdown version of [/jobs/ext/2756148-senior-data-engineer-identity-resolution-large-scale-data-engineering](https://www.wearedevelopers.com/jobs/ext/2756148-senior-data-engineer-identity-resolution-large-scale-data-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Engineer - Identity Resolution & Large-Scale Data Engineering - **Company:** Amazon.com, Inc. - **Location:** Bellevue, WA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Microsoft Azure, Big Data, Customer Data Management, Data Deduplication, Information Engineering, Data Systems, Distributed Computing Environment, Python (Programming Language), Machine Learning, Operational Databases, Performance Tuning, Standard Sql, Web Applications, Azure Data Factory, Large Language Models, Apache Spark, Generative AI, Scalability Testing, Pyspark, Low Latency, Apache Kafka, Cosmos DB, Spark Streaming, Virtual Agents, Data Pipelines, Serverless Computing, User Identification, Databricks - **Published:** September 6, 2026 - **Apply:** https://www.careerjet.com/jobad/us0af660969a254d7aedd8a1d7ff9385b9 ## About the Role The ideal candidate will have proven experience delivering data solutions at billion-record scale, with strong expertise in Azure Databricks, PySpark, Python, Azure Data Factory, and distributed data processing. The role requires strong production-scale engineering judgment-the ability to identify technical and scalability constraints early, evaluate alternatives quickly, and deliver tested, scalable, and production-ready solutions. Experience with Identity Resolution or comparable high-throughput matching systems is highly preferred., 7+ years of Data Engineering experience. Proven experience working with very large-scale datasets, preferably 1B+ records. Strong hands-on experience with Azure Databricks and PySpark. Strong Python and SQL skills. Experience with Azure Data Factory (ADF) and Azure data services. Strong understanding of distributed processing and Spark performance optimization. Experience delivering production-grade, scalable data solutions. Ability to identify technical constraints early and make timely technical decisions. Strong problem-solving and communication skills. Preferred Qualifications Experience with Identity Resolution / Entity Resolution / Record Linkage / Deduplication / Fuzzy Matching. Experience with Cosmos DB and low-latency APIs. Experience with Azure Functions or Azure Web Apps. Experience with Event Hub, Kafka, or Spark Structured Streaming. ML experience related to matching or customer data. Exposure to Generative AI, LLMs, RAG, or Agentic AI. Telecommunications or large-scale customer data experience. ## Description We are seeking an experienced Senior Data Engineer to design, build, and optimize production-grade data solutions supporting Identity Resolution (IDR) and large-scale customer data initiatives., Design and optimize "large-scale data processing solutions capable of operating at 1B+ record scale". Develop scalable data pipelines using Azure Databricks, PySpark, Python, and ADF". Optimize Spark workloads including large-scale joins, partitioning, shuffling, data skew, memory, and compute utilization. Design and improve solutions for Identity Resolution, Entity Resolution, Record Linkage, Deduplication, or similar high-volume matching problems. Evaluate technical approaches against production data volumes and infrastructure constraints. Identify and communicate scalability and environment constraints early and recommend viable alternatives. Perform performance and scalability testing using production-like workloads. Build reliable data pipelines with appropriate validation, error handling, monitoring, and recovery mechanisms. Collaborate with architects, engineers, data scientists, and stakeholders to deliver solutions from design through production. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Developer Tools for Microsoft Azure](https://www.wearedevelopers.com/videos/450-developer-tools-for-microsoft-azure) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production)