> Markdown version of [/jobs/ext/3538553-senior-data-engineer-applied-ai-solutions](https://www.wearedevelopers.com/jobs/ext/3538553-senior-data-engineer-applied-ai-solutions). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Engineer, Applied AI Solutions - **Company:** Amazon.com, Inc. - **Location:** Seattle, WA, United States - **Experience:** Expert - **Salary:** $154,600.0 - $209,100.0 - **Contract:** Permanent contract - **Skills:** Sql Data Warehouse, Java (Programming Language), Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Data Analysis, Big Data, Databases, Information Engineering, Data Infrastructure, Extract Transform Load (ETL), Data Mining, Data Systems, Distributed Systems, Apache Hadoop, Apache Hive, Python (Programming Language), Node.Js, Performance Tuning, Standard Sql, Scala (Programming Language), Data Streaming, Workflow Management Systems, Scripting, Data Storage Technologies, Sql Optimization, Retrieval-Augmented Generation, Multi-Agent Systems, Apache Spark, Electronic Medical Records, Information Technology, Data Lineage, Amazon Redshift, Programming Languages - **Published:** October 1, 2026 - **Apply:** https://dejobs.org/x/x/6EE493C6697E48EF896B00DBE427FE5A/job/ ## About the Role * Experience building data infrastructure that serves AI systems and autonomous agents, not just human analysts, including machine-consumable interfaces, metadata, and documentation. * Experience with GenAI data patterns end to end: chunking, embeddings, and vector stores for retrieval-augmented generation. * Experience building and maintaining datasets and feature pipelines for ML/GenAI training, fine-tuning, and inference (Amazon SageMaker, Bedrock, or equivalent). * Experience implementing data quality, lineage, and observability that AI workloads depend on including validation, freshness/anomaly monitoring, and alerting at scale. * 5+ years of Python (or Scala/Java) and advanced SQL, including performance tuning at scale. * Experience with batch and streaming ETL/ELT on AWS (Glue, EMR/Spark, S3, Athena) and a production cloud data warehouse (Amazon Redshift or equivalent). * Experience designing data models and schemas for analytical, operational, and AI/retrieval workloads. * Experience with workflow orchestration (Step Functions, Airflow, or Glue Workflows)., * 7+ years of data engineering experience * Experience with data modeling, warehousing and building ETL pipelines * Experience with SQL * Experience in at least one modern scripting or programming language, such as Python, Java, Scala, or NodeJS * Experience mentoring team members on best practices * Experience with MPP databases such as Amazon Redshift * Experience building/operating highly available, distributed systems of data extraction, ingestion, and processing of large data sets * Experience building data infrastructure that serves AI systems and autonomous agents, not just human analysts, including machine-consumable interfaces, metadata, and documentation. * Experience with GenAI data patterns end to end: chunking, embeddings, and vector stores for retrieval-augmented generation. * Experience building and maintaining datasets and feature pipelines for ML/GenAI training, fine-tuning, and inference (Amazon SageMaker, Bedrock, or equivalent). * Experience implementing data quality, lineage, and observability that AI workloads depend on including validation, freshness/anomaly monitoring, and alerting at scale., * Experience with big data technologies such as: Hadoop, Hive, Spark, EMR * Experience operating large data warehouses * Experience providing technical leadership and mentoring other engineers for best practices on data engineering * Bachelor's degree in computer science, engineering, analytics, mathematics, statistics, IT or equivalent * Knowledge of distributed systems as it pertains to data storage and computing ## Description The newest business group in AWS, Applied AI Solutions are built by AWS and AWS Partners to deliver applied AI solutions that leverage Amazon's operational expertise and that businesses love and trust for their day-to-day success. Our ambition is to become a partner which companies can rely on to run their business every day, putting AI to work delivering better customer experience, operational excellence and speed. We are seeking a Senior Data Engineer to design, build and maintain our next-generation data infrastructure - one that seamlessly serves both human analysts and AI systems. This role sits at the intersection of traditional enterprise data warehousing and innovative AI technologies, requiring someone who can bridge these worlds to create a unified, future-proof data ecosystem. As a key member of our data team, you'll collaborate across organizational boundaries with data scientists, engineers, analytics teams, and business stakeholders to develop innovative and scalable solutions that push the boundaries of what's possible with our data assets. You'll be responsible for ensuring our datasets maintain the highest levels of accuracy, consistency, and observability - implementing comprehensive monitoring, lineage tracking, and self-healing mechanisms that maintain data quality at scale. Your infrastructure will support both analysts / scientists and autonomous AI agents with equal effectiveness, requiring thoughtful interfaces, documentation, and metadata that serve both audiences. In this role, you'll champion a forward-thinking approach to data infrastructure that anticipates the evolving needs of AI systems while maintaining the reliability and performance that business operations demand. You'll help shape our technical roadmap for data systems that will serve as the foundation for our organization's AI transformation journey., * 5+ years of data engineering, building and operating production pipelines and warehouses. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Fully Orchestrating Databricks from Airflow](https://www.wearedevelopers.com/videos/336-fully-orchestrating-databricks-from-airflow) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)