> Markdown version of [/jobs/ext/205011-support-engineer-aws-incident-response](https://www.wearedevelopers.com/jobs/ext/205011-support-engineer-aws-incident-response). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Support Engineer, AWS Incident Response - **Company:** Amazon.com, Inc. - **Location:** Seattle, WA, United States - **Experience:** Experienced - **Salary:** $90,400.0 - $158,200.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Data Analysis, Bash Shell, Cloud Computing, Computer Programming, Linux, Distributed Systems, Python (Programming Language), Networking Basics, Operational Data Store, Datadog, Scripting, Grafana, Generative AI, Cloudwatch, Golang - **Published:** May 19, 2026 - **Apply:** https://www.jobmonkeyjobs.com/career/27690789/Support-Engineer-Aws-Incident-Response-Washington-Seattle-7375 ## About the Role 2+ years of technical support experience - Direct experience participating in incident response for production systems - Strong understanding of operating systems (Linux), networking fundamentals, and distributed systems - Experience with operational monitoring, alerting, and metrics (CloudWatch, Datadog, Grafana, or equivalent) - Demonstrated ability to troubleshoot complex technical problems spanning multiple systems or services - Experience scripting or programming in at least one modern language (Python, Bash, Go, or similar) - Ability to clearly break down technical complexity for a wide range of audiences, from engineers to senior leadership, without relying on jargon, Familiarity with incident management tooling and workflows - Experience with AWS services and cloud infrastructure - Experience using generative AI or automation to solve operational problems or accelerate workflows - Track record of authoring post-incident analyses (post-mortems) and driving corrective actions to completion - Experience building operational dashboards, runbooks, or automation that improved team efficiency - Experience coordinating across globally distributed teams and time zones ## Description AWS Incident Response (AIR) keeps AWS working for millions of customers. When major incidents hit, AIR leads the response, coordinating resolvers across AWS and driving mitigation. We move fast, but not carelessly, obsessing over observability of the cloud and perpetually improving our detection and response speed and accuracy. We ensure each incident drives improvements that strengthen AWS. It's a high-visibility, high-impact role with a global view of AWS health that few teams get to see. The Role As a Support Engineer on AIR's Seattle team, you'll be on the front line of AWS incident response. You'll lead high-severity calls, triage complex failures across distributed systems, coordinate resolver teams, and drive incidents to mitigation while millions of customers depend on the outcome. Between incidents, you'll obsess over metrics and detection analysis, building dashboards and mechanisms that surface problems before customers notice. You will drive operational improvements that make the incident management ecosystem faster and more accurate. This isn't a role where you watch dashboards and robotically follow runbooks. You'll deep-dive the largest, most complex technical environment in the world. You'll develop expertise across AWS services, networking, and infrastructure. You'll own operational processes end-to-end and use data to find the next leap in how we protect the cloud. If interested, you'll also have the opportunity to grow your development skills by taking on coding projects matched to your ability level. This role includes participation in an on-call rotation, including some weekends and holidays. Key job responsibilities Incident Response Lead high-severity incident response calls. Triage, coordinate resolvers across AWS service teams, communicate clearly under pressure, and drive incidents to mitigation. Manage escalations and ensure accurate documentation throughout. Operational Excellence and Detection Own and run operational health reviews. Build and maintain dashboards, metrics, and monitoring that surface trends before they become incidents. Obsess over detection accuracy and speed. Detect patterns across events and drive proactive mechanisms to prevent recurrence. Metrics and Analysis Deep-dive operational data to identify systemic issues, measure response effectiveness, and prioritize improvements. Use metrics to tell the story of what's working, what's degrading, and where the next risk is hiding. Process and Tooling Improvement Identify gaps in operational processes, documentation, and tooling. Build or improve mechanisms that reduce time-to-detection and time-to-mitigation. Use data to prioritize where effort has the highest impact. Automation and Generative AI Leverage scripting, generative AI, and automation to accelerate incident response, improve detection, and reduce toil. Identify opportunities where AI can augment human judgment during incidents or surface insights from operational data at scale. Driving Continuous Improvement Ensure each incident makes AWS stronger. Work with service teams to ensure learnings from incidents drive corrective actions and that follow-through happens. Close the loop between what broke and what gets fixed. ## Related Videos - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Dev Digest 162: AI careers, MCP, AWS best practices & floppy sweaters](https://www.wearedevelopers.com/magazine/571-dev-digest-162-ai-careers-mcp-aws-best-practices-floppy-sweaters) - [How to Answer the Interview Question: “Why Do You Want to Be a Software Engineer?”](https://www.wearedevelopers.com/magazine/392-how-to-answer-the-interview-question-why-do-you-want-to-be-a-software-engineer) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated)