Infrastructure Reliability Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+10 more
Job description
Amazon Web Services is seeking a Sr. Infrastructure Reliability Engineer to build and operate highly reliable, scalable, and secure cloud infrastructure. You will design and automate provisioning, monitoring, and recovery for large-scale services, using IaC, CI/CD, and observability tools to prevent and resolve incidents. Partnering with software, security, and operations teams, you’ll drive root-cause analysis, improve availability and performance, and champion operational excellence in a fast-paced, customer-obsessed environment while learning cutting-edge AWS technologies.
Responsibilities
- Design and implement highly reliable, scalable AWS infrastructure for production services.
- Automate provisioning, configuration, and deployments using Infrastructure as Code and CI/CD.
- Develop and maintain monitoring, alerting, and observability dashboards for critical systems.
- Lead and participate in on-call rotation, incident response, and post-incident reviews.
- Drive root-cause analysis and implement long-term fixes to improve availability and resiliency.
- Collaborate with software, security, and operations teams to enforce best practices and standards.
- Optimize performance, capacity, and cost across infrastructure components.
- Enhance reliability through chaos testing, fault injection, and resilience patterns.
- Document runbooks, operational procedures, and architectures for supported services.
- Mentor engineers on reliability engineering, automation, and AWS best practices.
Requirements
- AWS cloud services (EC2, S3, RDS, VPC)
- Infrastructure as Code (Terraform/Cloud
- Formation)
- Linux systems administration
- CI/CD pipelines (e.g., Code
- Pipeline, Jenkins)
- Monitoring & observability (Cloud
- Watch, Prometheus, Grafana)
- Scripting (Python, Bash, or similar)
- Containerization & Kubernetes
- Networking & security fundamentals
- Incident management & on-call support
- Performance tuning & capacity planning
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What Are The Top Skills Required For Azure Developers?
Highest Paying Tech Companies for Developers
7 Cloud Computing Trends Coming in 2025 for Developers
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence