> Markdown version of [/jobs/ext/2709433-data-engineer](https://www.wearedevelopers.com/jobs/ext/2709433-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Amazon.com, Inc. - **Location:** Seattle, WA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Amazon Web Services, Apache HTTP Server, Big Data, Computer Programming, Continuous Delivery, Continuous Integration, Information Engineering, Data Governance, Data Integrity, Extract Transform Load (ETL), Data Transformation, Data Migration, Data Mining, Data Stores, Distributed Data Store, Apache Hadoop, JSON, Job Scheduling, Python (Programming Language), Performance Tuning, Systems Development Life Cycle, Standard Sql, Software Engineering, SQL Databases, Parquet, Data Processing, File Transfer Protocol (FTP), Cloud Platform System, Snowflake, Apache Spark, Sybase, Build Management, Data Lakes, Ansi Sql, Semi-structured Data, Core Data, Kubernetes, Avro, Apache Kafka, Data Management, Code Restructuring, Software Version Control, Data Pipelines - **Published:** September 4, 2026 - **Apply:** https://www.careerjet.com/jobad/us12ef718632520178260ec310ae3fb507 ## About the Role * Minimum 3+ years of hands-on software development and/or data engineering experience, including coding, development, troubleshooting, and implementation of data-related solutions. * Minimum 3+ years of experience working with SQL, including development, analysis, query troubleshooting, and performance optimization. * Minimum 3+ years of hands-on programming experience using Python and/or Java for data processing, application development, automation, or integration. * Minimum 3+ years of experience developing or supporting ETL/ELT data pipelines, including data extraction, transformation, loading, and pipeline troubleshooting. * Experience developing or supporting distributed data-processing solutions using Apache Spark. * Experience working with data engineering concepts including SCD Type 2, schema evolution, partitioning, clustering, normalization versus denormalization, natural versus surrogate keys, and data quality frameworks. * Experience working with one or more data and integration technologies including Kafka, ANSI SQL, FTP, Apache Spark, Hadoop, Snowflake, Apache Iceberg, and Sybase IQ. * Experience working with structured and semi-structured data formats including JSON, Avro, and Parquet. * Experience working within established SDLC and CI/CD processes, including source control, testing, deployment, and release practices. * Familiarity with containerized application environments and Kubernetes. * Demonstrated ability to troubleshoot technical issues, communicate effectively with stakeholders, collaborate across global teams, and take ownership of assigned deliverables. Nice to Have: * Experience performing large-scale data platform or datastore migration initiatives. * Experience migrating data workloads from on-premises environments to cloud-based platforms, particularly AWS. * Experience working with Lakehouse architectures. * Experience with Snowflake and Apache Iceberg-based data platforms. * Experience working within the financial services industry. * Experience supporting complex migration programs involving multiple business, technology, and global delivery teams. ## Description * Perform end-to-end datastore migrations from on-premises DataLake environments to AWS-hosted Lakehouse platforms as part of the migration factory team. * Refactor and migrate existing data pipelines, including data extraction logic, transformation processes, and job scheduling. * Execute large-scale data transfers while ensuring data integrity, completeness, reliability, and consistency between source and target platforms. * Translate and modernize legacy SQL and Apache Spark-based processing logic for Snowflake and Apache Iceberg environments. * Analyze existing data usage patterns, business requirements, and downstream consumption to support the development and delivery of reusable data products. * Design and build data reconciliation and validation frameworks to verify data accuracy during and after migration. * Collaborate with business stakeholders, application teams, data owners, and technical teams to perform validation and obtain migration sign-off. * Act as a technical liaison between migration, data engineering, application, infrastructure, and business teams throughout the migration lifecycle. * Troubleshoot data pipeline, data transformation, performance, and migration-related issues and implement appropriate technical solutions. * Optimize data pipelines and processing workloads to improve performance, scalability, and operational efficiency. * Apply core data engineering concepts including Slowly Changing Dimensions (SCD Type 2), schema evolution, partitioning, clustering, normalization and denormalization, natural and surrogate keys, and data quality controls. * Work with structured and semi-structured data formats including JSON, Avro, and Parquet. * Follow established Software Development Life Cycle (SDLC), source-control, testing, deployment, and Continuous Integration/Continuous Deployment (CI/CD) practices. * Adapt to new data technologies, migration tools, engineering standards, and workflows as required by the program. * Collaborate effectively with geographically distributed and global delivery teams. ## Related Videos - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [From event streaming to event sourcing 101](https://www.wearedevelopers.com/videos/91-from-event-streaming-to-event-sourcing-101) - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [Introducing JSON Structure](https://www.wearedevelopers.com/videos/100219-introducing-json-structure) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Green Cloud Computing](https://www.wearedevelopers.com/videos/592-green-cloud-computing) ## Related Articles - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers)