Software Development Engineer, Infrastructure Reliability Engineering

Amazon.com, Inc.
Nashville, TN, United States
2 days ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Internship / Graduate position
Employment type
Full-time (> 32 hours)
Experience required
1 year minimum
Compensation
$136,500.0 - $184,700.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Artificial Intelligence Amazon Web Services Software Applications C Sharp (Programming Language) C++ (Programming Language) Software Quality Code Review Computer Programming Software Design Patterns Perl (Programming Language) Machine Learning
+11 more
Object-Oriented Software Development Reliability Engineering Software Engineering Multithreading Mobile Robots Large Language Models Information Technology Build Process Software Coding Software Version Control Serverless Computing

Job description

Join us in building Reflex, the agentic incident-response platform for Amazon’s fulfillment network. You’ll design and ship production software and AI agents on Amazon Bedrock AgentCore that triage high-severity incidents, generate real-time call intelligence, draft stakeholder communications, and produce structured post-incident records, shifting incident response from a manual, pull-based model to an intelligent, push-based one.

Amazon’s network of fulfillment centers is the infrastructure that Amazon Robotics runs on. When it degrades, robots stop and packages stop moving. The team manages thousands of high-severity incidents every year, with Incident Managers assembling context across many systems under time pressure before resolution work can even begin. The software you build puts that context in front of responders in seconds and removes the repetitive work that extends incidents today. You’ll design the data, evaluation, and feedback mechanisms that make the system measurably better with every incident it touches.

This role sits on a software team within Operations Infrastructure Services (OIS), part of Amazon Robotics. You’ll stay close to live operations through incident reviews, workflow observation, and call shadowing, and turn what you learn into durable software. Reflex is the immediate focus; as it matures, the same foundations are expected to extend toward shared incident context across organizations, coordinated agent workflows, and carefully guarded automation of recovery validation and repeatable response actions, with every step gated by measurable confidence., * Design, build, test, deploy, and operate production services and AI agents on AWS (Amazon Bedrock AgentCore, serverless compute, event-driven pipelines) that automate incident triage, call intelligence, communications, and post-incident documentation and reporting

  • Own features end-to-end: from discovery with Incident Managers and resolver teams, through design, implementation, evaluation, deployment, and production operation
  • Build the foundations that gate agent autonomy: LLM output evaluation, observability and alerting for agents in production, and identity and access controls aligned with Amazon standards
  • Design the feedback loops through which agents learn: capturing human reviews, corrections, approvals, and incident outcomes as evaluation signal, and turning resolved incidents into structured history that improves recommendations over time
  • Integrate with the incident lifecycle (ticketing, chat, telemetry, detection feeds, and live call transcription) and model consistent incident state across those systems
  • Raise the bar on software quality, security, testing, and operational excellence for AI systems acting inside production incident workflows

A day in the life

You start by reviewing overnight agent evaluation results and fixing a class of inaccurate triage recommendations, shipping an improvement responders see on the next incident. Later, you pair with an Incident Manager to see how they corrected an agent-drafted call summary, then turn that correction into an automated evaluation case and an agent fix. You might close the day reviewing a design for representing incident state across source systems, or shadowing part of a live bridge call to spot where responders lose time. The operational insight you gather today becomes the software that shortens tomorrow’s incident.

Requirements

  • 3+ years of non-internship professional software development experience
  • 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • 1+ years of software development engineer or related occupational experience
  • 1+ years of designing and developing large-scale, multi-tiered, multi-threaded, embedded or distributed software applications, tools, systems, and services using: C#, C++, Java, or Perl experience
  • 1+ years of Object Oriented Design experience
  • Bachelor’s degree or foreign equivalent in Computer Science, Engineering, Mathematics, or a related field
  • Experience programming with at least one software programming language, * 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques

Benefits & conditions

Amazon offers a full range of benefits to support you and eligible family members, including domestic partners and children. Benefits can vary by location, the number of regularly scheduled hours you work, length of employment, and job status such as seasonal or temporary employment. The benefits that generally apply to regular, full-time employees include:

  1. Medical, Dental, and Vision Coverage
  2. Maternity and Parental Leave Options
  3. Paid Time Off (PTO)
  4. 401(k) Plan, The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .

USA, TN, Nashville - 136,500.00 - 184,700.00 USD annually

USA, VA, Arlington - 143,700.00 - 194,400.00 USD annually

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

43 sec

Software engineering journey and local Manchester roots

Jonathan Tang · Coffee With Developers

3:39 min

Addressing code review surrender and process exploitation

Laura Tacho Laura Tacho · World Congress 2026 Europe

2:59 min

Generating a Swagger JSON file during the build process

Roman Alexis Anastasini · World Congress 2021

1:19 min

Evaluating headlines on chatbots, robot mobility, and AI appeal

Chris Heilmann +1 · LIVE

1:39 min

Speaker introduction and software engineering background

Daniel Raniz Raneland · LIVE

56 sec

The hidden costs of delayed peer code reviews

Tim Gilboy Tim Gilboy

Videos

See all

Related articles

See all