Senior Data Engineer, Applied AI Solutions

Amazon.com, Inc.
Seattle, WA, United States
3 days ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$154,600.0 - $209,100.0
Working hours
Regular working hours
Job source

Tech stack

Sql Data Warehouse Java (Programming Language) Artificial Intelligence Airflow Amazon Web Services Amazon S3 Data Analysis Big Data Databases Information Engineering Data Infrastructure Extract Transform Load (ETL)
+23 more
Data Mining Data Systems Distributed Systems Apache Hadoop Apache Hive Python (Programming Language) Node.Js Performance Tuning Standard Sql Scala (Programming Language) Data Streaming Workflow Management Systems Scripting Data Storage Technologies Sql Optimization Retrieval-Augmented Generation Multi-Agent Systems Apache Spark Electronic Medical Records Information Technology Data Lineage Amazon Redshift Programming Languages

Job description

The newest business group in AWS, Applied AI Solutions are built by AWS and AWS Partners to deliver applied AI solutions that leverage Amazon’s operational expertise and that businesses love and trust for their day-to-day success. Our ambition is to become a partner which companies can rely on to run their business every day, putting AI to work delivering better customer experience, operational excellence and speed.

We are seeking a Senior Data Engineer to design, build and maintain our next-generation data infrastructure - one that seamlessly serves both human analysts and AI systems. This role sits at the intersection of traditional enterprise data warehousing and innovative AI technologies, requiring someone who can bridge these worlds to create a unified, future-proof data ecosystem.

As a key member of our data team, you’ll collaborate across organizational boundaries with data scientists, engineers, analytics teams, and business stakeholders to develop innovative and scalable solutions that push the boundaries of what’s possible with our data assets.

You’ll be responsible for ensuring our datasets maintain the highest levels of accuracy, consistency, and observability - implementing comprehensive monitoring, lineage tracking, and self-healing mechanisms that maintain data quality at scale. Your infrastructure will support both analysts / scientists and autonomous AI agents with equal effectiveness, requiring thoughtful interfaces, documentation, and metadata that serve both audiences.

In this role, you’ll champion a forward-thinking approach to data infrastructure that anticipates the evolving needs of AI systems while maintaining the reliability and performance that business operations demand. You’ll help shape our technical roadmap for data systems that will serve as the foundation for our organization’s AI transformation journey., * 5+ years of data engineering, building and operating production pipelines and warehouses.

Requirements

  • Experience building data infrastructure that serves AI systems and autonomous agents, not just human analysts, including machine-consumable interfaces, metadata, and documentation.
  • Experience with GenAI data patterns end to end: chunking, embeddings, and vector stores for retrieval-augmented generation.
  • Experience building and maintaining datasets and feature pipelines for ML/GenAI training, fine-tuning, and inference (Amazon SageMaker, Bedrock, or equivalent).
  • Experience implementing data quality, lineage, and observability that AI workloads depend on including validation, freshness/anomaly monitoring, and alerting at scale.
  • 5+ years of Python (or Scala/Java) and advanced SQL, including performance tuning at scale.
  • Experience with batch and streaming ETL/ELT on AWS (Glue, EMR/Spark, S3, Athena) and a production cloud data warehouse (Amazon Redshift or equivalent).
  • Experience designing data models and schemas for analytical, operational, and AI/retrieval workloads.
  • Experience with workflow orchestration (Step Functions, Airflow, or Glue Workflows)., * 7+ years of data engineering experience
  • Experience with data modeling, warehousing and building ETL pipelines
  • Experience with SQL
  • Experience in at least one modern scripting or programming language, such as Python, Java, Scala, or NodeJS
  • Experience mentoring team members on best practices
  • Experience with MPP databases such as Amazon Redshift
  • Experience building/operating highly available, distributed systems of data extraction, ingestion, and processing of large data sets
  • Experience building data infrastructure that serves AI systems and autonomous agents, not just human analysts, including machine-consumable interfaces, metadata, and documentation.
  • Experience with GenAI data patterns end to end: chunking, embeddings, and vector stores for retrieval-augmented generation.
  • Experience building and maintaining datasets and feature pipelines for ML/GenAI training, fine-tuning, and inference (Amazon SageMaker, Bedrock, or equivalent).
  • Experience implementing data quality, lineage, and observability that AI workloads depend on including validation, freshness/anomaly monitoring, and alerting at scale., * Experience with big data technologies such as: Hadoop, Hive, Spark, EMR
  • Experience operating large data warehouses
  • Experience providing technical leadership and mentoring other engineers for best practices on data engineering
  • Bachelor’s degree in computer science, engineering, analytics, mathematics, statistics, IT or equivalent
  • Knowledge of distributed systems as it pertains to data storage and computing

Benefits & conditions

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .

USA, WA, Seattle - 154,600.00 - 209,100.00 USD annually

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset · World Congress 2023

2:19 min

Introduction to Apache Airflow for advanced orchestration

Alan Mazankiewicz · LIVE

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

1:59 min

Evolving roles in AI driven software teams

Ignacio Riesgo Ignacio Riesgo +1 · World Congress 2024

3:55 min

Identifying underlying Node.js runtime vulnerabilities using fuzzing tools

Sonya Moisset · World Congress 2023

Videos

See all

Related articles

See all