Infrastructure Reliability Engineer

Amazon.com, Inc.
Herndon, VA, United States
16 days ago
Apply on find.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$145,000.0 - $190,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Bash Shell Cloud Computing Continuous Integration Monitoring of Systems Python (Programming Language) Network Security Linux System Administration Performance Tuning Reliability Engineering
+10 more
Prometheus Scripting Grafana Amazon Virtual Private Cloud (VPC) Amazon Relational Database Service Containerization Kubernetes Deployment Automation Terraform Jenkins

Job description

Amazon Web Services is seeking a Sr. Infrastructure Reliability Engineer to build and operate highly reliable, scalable, and secure cloud infrastructure. You will design and automate provisioning, monitoring, and recovery for large-scale services, using IaC, CI/CD, and observability tools to prevent and resolve incidents. Partnering with software, security, and operations teams, you’ll drive root-cause analysis, improve availability and performance, and champion operational excellence in a fast-paced, customer-obsessed environment while learning cutting-edge AWS technologies.

Responsibilities

  • Design and implement highly reliable, scalable AWS infrastructure for production services.
  • Automate provisioning, configuration, and deployments using Infrastructure as Code and CI/CD.
  • Develop and maintain monitoring, alerting, and observability dashboards for critical systems.
  • Lead and participate in on-call rotation, incident response, and post-incident reviews.
  • Drive root-cause analysis and implement long-term fixes to improve availability and resiliency.
  • Collaborate with software, security, and operations teams to enforce best practices and standards.
  • Optimize performance, capacity, and cost across infrastructure components.
  • Enhance reliability through chaos testing, fault injection, and resilience patterns.
  • Document runbooks, operational procedures, and architectures for supported services.
  • Mentor engineers on reliability engineering, automation, and AWS best practices.

Requirements

  • AWS cloud services (EC2, S3, RDS, VPC)
  • Infrastructure as Code (Terraform/Cloud
  • Formation)
  • Linux systems administration
  • CI/CD pipelines (e.g., Code
  • Pipeline, Jenkins)
  • Monitoring & observability (Cloud
  • Watch, Prometheus, Grafana)
  • Scripting (Python, Bash, or similar)
  • Containerization & Kubernetes
  • Networking & security fundamentals
  • Incident management & on-call support
  • Performance tuning & capacity planning

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on find.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:31 min

Implementing initial GitOps architecture with Jenkins and Argo CD

Lian Li · World Congress 2022

1:07 min

Architecting the availability stack with Prometheus and Grafana

Gabriel Labachelerie · World Congress 2023

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

4:23 min

Reviewing AWS infrastructure deployment configuration and planning

Devlin Duldulao · LIVE

2:38 min

Challenges with legacy Jenkins and Groovy pipelines

Martin Beránek · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all