Senior Data Engineer, Infrastructure Reliability

Amazon.com, Inc.
Austin, TX, United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$154,600.0 - $209,100.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Artificial Intelligence Data Analysis Batch Processing Big Data Software Quality Information Engineering Extract Transform Load (ETL) Data Retention Amazon DynamoDB Apache Hadoop Apache Hive
+16 more
Python (Programming Language) Machine Learning Node.Js Operational Databases Standard Sql DataOps Scala (Programming Language) Scripting Feature Engineering Data Ingestion Apache Spark Electronic Medical Records Data Lakes Information Technology Cloudwatch Programming Languages

Job description

Help build the data foundation that keeps Amazon’s fulfillment network running 24/7. Infrastructure Reliability is building an AI-powered platform that detects, diagnoses, and resolves incidents across thousands of sites globally, and none of it works without clean, reliable, well-modeled data.

You will own and evolve production data pipelines that power real-time incident detection and correlation, build the data lake and unified data models that unblock ML and autonomous resolution initiatives, and rescue historical telemetry data before it expires and becomes permanently unrecoverable. This is a high-ownership role with direct, visible impact on Amazon’s global fulfillment operations, working closely with applied scientists who train models on the data you build., You will design, build, and operate ETL pipelines that ingest, transform, and correlate incident and telemetry data from sources including DynamoDB, OpenSearch, CloudWatch, and internal incident management systems, publishing curated datasets to our data lake and Andes for consumption by dashboards, applications, and ML models. You will take ownership of production pipelines, including scheduled batch processing jobs and Glue-based ETL workflows, ensuring reliability, monitoring, and timely resolution of pipeline issues.

You will build data retention and ingestion pipelines to preserve time-bound telemetry signals before source-system expiration windows, and you will design standardized ingestion frameworks that normalize data from many disparate sources into a common, science-consumable format. You will contribute to the design of a unified data lake architecture, including schema design, partitioning strategy, and access patterns, replacing fragmented, duplicated pipelines with a single source of truth.

You will partner closely with applied scientists and engineers to grasp data requirements for ML model training and feature engineering, and you will mentor other engineers on data engineering best practices, code quality, and pipeline design.

A day in the life You might start your day investigating a pipeline failure alert, tracing it back through Glue job logs to a schema change upstream. Later, you’re in a design discussion with an applied scientist about what shape of data would best enable a new model, translating that into a concrete schema. In the afternoon, you’re heads-down building a new ingestion adapter or reviewing a teammate’s pull request. Your work directly determines what data is available, and reliable, for the platform’s detection and reasoning capabilities.

Requirements

  • 5+ years of data engineering experience
  • Experience with data modeling, warehousing and building ETL pipelines
  • Experience with SQL
  • Experience in at least one modern scripting or programming language, such as Python, Java, Scala, or NodeJS
  • Experience mentoring team members on best practices
  • Bachelor’s degree in computer science, engineering, analytics, mathematics, statistics, IT or equivalent, * Experience with big data technologies such as: Hadoop, Hive, Spark, EMR
  • Experience operating large data warehouses
  • Master’s degree in computer science, engineering, analytics, mathematics, statistics, IT or equivalent

Benefits & conditions

AD&D insurance, Parental leave, 401(k), Health insurance, 401(k) matching, Paid time off, Vision insurance, Dental insurance Full-time Austin, TX, Amazon offers a full range of benefits that support you and eligible family members, including domestic partners. Benefits can vary by location, the number of regularly scheduled hours you work, length of employment, and job status such as seasonal or temporary employment. The benefits that generally apply to regular, full-time employees include:

  1. Medical, Dental, and Vision Coverage
  2. Maternity and Parental Leave Options
  3. Paid Time Off (PTO)
  4. 401(k) Plan

If you are not sure that every qualification on the list above describes you exactly, we’d still love to hear from you! At Amazon, we value people with unique backgrounds, experiences, and skillsets. If you’re passionate about this role and want to make an impact on a global scale, please apply!

About the team Infrastructure Reliability sits within Amazon’s Robotics organization, building the platform that keeps fulfillment operations running no matter what breaks. We do not own any single domain; we build the data and orchestration layer that sees across all of them, identifying failures that cascade across team boundaries. We are a small, technically deep team building AI-powered detection and remediation capabilities at scale. We value ownership, rigor, and hands-on technical depth, and we move quickly from idea to production.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset · World Congress 2023

2:00 min

Separating dataset creation from low-level software implementation steps

Jan Zawadzki · World Congress 2022

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

3:55 min

Identifying underlying Node.js runtime vulnerabilities using fuzzing tools

Sonya Moisset · World Congress 2023

Videos

See all

Related articles

See all