> Markdown version of [/jobs/ext/2942002-infrastructure-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2942002-infrastructure-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Infrastructure Reliability Engineer - **Company:** Amazon.com, Inc. - **Location:** Herndon, VA, United States - **Experience:** Expert - **Salary:** $145,000.0 - $190,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Bash Shell, Cloud Computing, Continuous Integration, Monitoring of Systems, Python (Programming Language), Network Security, Linux System Administration, Performance Tuning, Reliability Engineering, Prometheus, Scripting, Grafana, Amazon Virtual Private Cloud (VPC), Amazon Relational Database Service, Containerization, Kubernetes, Deployment Automation, Terraform, Jenkins - **Published:** September 16, 2026 - **Apply:** https://find.jobs/jobs-near-me/apply/ats-redirect/?id=2974387148-2 ## About the Role * AWS cloud services (EC2, S3, RDS, VPC) * Infrastructure as Code (Terraform/Cloud * Formation) * Linux systems administration * CI/CD pipelines (e.g., Code * Pipeline, Jenkins) * Monitoring & observability (Cloud * Watch, Prometheus, Grafana) * Scripting (Python, Bash, or similar) * Containerization & Kubernetes * Networking & security fundamentals * Incident management & on-call support * Performance tuning & capacity planning ## Description Amazon Web Services is seeking a Sr. Infrastructure Reliability Engineer to build and operate highly reliable, scalable, and secure cloud infrastructure. You will design and automate provisioning, monitoring, and recovery for large-scale services, using IaC, CI/CD, and observability tools to prevent and resolve incidents. Partnering with software, security, and operations teams, you'll drive root-cause analysis, improve availability and performance, and champion operational excellence in a fast-paced, customer-obsessed environment while learning cutting-edge AWS technologies. Responsibilities * Design and implement highly reliable, scalable AWS infrastructure for production services. * Automate provisioning, configuration, and deployments using Infrastructure as Code and CI/CD. * Develop and maintain monitoring, alerting, and observability dashboards for critical systems. * Lead and participate in on-call rotation, incident response, and post-incident reviews. * Drive root-cause analysis and implement long-term fixes to improve availability and resiliency. * Collaborate with software, security, and operations teams to enforce best practices and standards. * Optimize performance, capacity, and cost across infrastructure components. * Enhance reliability through chaos testing, fault injection, and resilience patterns. * Document runbooks, operational procedures, and architectures for supported services. * Mentor engineers on reliability engineering, automation, and AWS best practices. ## Related Videos - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [Monitoring as Code - Managing your dashboards at scale](https://www.wearedevelopers.com/videos/753-monitoring-as-code-managing-your-dashboards-at-scale) - [Reliable scalability: How Amazon.com scales on AWS](https://www.wearedevelopers.com/videos/983-reliable-scalability-how-amazon-com-scales-on-aws) - [GitLab CI pipelines for a whole company](https://www.wearedevelopers.com/videos/143-gitlab-ci-pipelines-for-a-whole-company) - [Terraform for Developers](https://www.wearedevelopers.com/videos/3-terraform-for-developers) - [Our GitOps approach for deploying an Identity Provider and an API Gateway in a SaaS company](https://www.wearedevelopers.com/videos/776-our-gitops-approach-for-deploying-an-identity-provider-and-an-api-gateway-in-a-saas-company) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)