Software Development Manager, EC2 UltraServer Availability

Amazon.com, Inc.
Seattle, WA, United States
29 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$184,900.0 - $250,200.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Amazon Elastic Compute Cloud Code Review Computer Engineering Data Centers Software Design Patterns Software Product Management Software Engineering Supercomputing Web Services Build Process
+3 more
Machine Learning Operations Software Coding Software Version Control

Job description

The EC2 UltraServer Availability team builds the automated repair and recovery systems for Amazon’s GB200 and GB300 fleet. One of the fastest-growing and most visible infrastructure domains at AWS. This is early-days territory: many repair flows that are fully automated for traditional EC2 hosts are still being invented for UltraServers. You’ll set the technical direction that turns manual, expert-driven recovery into deterministic, self-healing automation, and you’ll see your team’s impact directly in fleet availability numbers that leadership watches weekly., You’ll lead a two-pizza team of 8-12 engineers automating end-to-end repair of entire UltraServers: detection, control-plane teardown, chassis swap orchestration with hardware engineering and data center operations, fabric rebuild, burn-in testing, and return to customer.

Your north-star metrics are fleet availability and repair dwell time: how fast a broken UltraServer gets back to serving customers.

Examples of projects your team is working on

  1. Automating deterministic repair: building the system that maps a failure signature directly to the right repair action, removing humans from the loop for well-understood failures

  2. Orchestrating multi-team repair procedures that today require coordination across hardware engineering, data center technicians, and EC2 control-plane services

  3. Cutting repair dwell time by parallelizing teardown, physical repair, and validation steps that currently run serially

  4. Extending recovery workflows into new regions, including air-gapped (ADC) environments

A day in the life You’ll spend most of your time where an SDM should: growing engineers (1:1s, career development, hiring), making roadmap and architecture calls with your senior engineers, and unblocking the team. Because we sit at the intersection of software, hardware, and data center operations, you’ll regularly partner with EC2 ML Supercomputing, hardware engineering, and DCO leadership. You’ll also drive operational reviews - this is a fleet-facing team with an on-call rotation, and reducing that operational load through automation is itself a core part of the charter. You’ll communicate progress and strategy to senior leadership through narratives and business reviews.

About the team The EC2 UltraServer Availability team (part of EC2 Nitro) maintains the availability of NVIDIA-based ML infrastructure at scale. We own end-to-end recovery and repair for GB200 and GB300 UltraServers: from detecting an availability event through repair, testing, and return to service. We work closely with hardware engineering, data center operations, EC2 ML Supercomputing, and EC2 capacity teams. We value engineers who are curious about the full stack, from NVLink cabling to control-plane APIs, and we invest in reducing our own operational toil through automation.

Requirements

  • 3+ years of engineering team management experience
  • 7+ years of working directly within engineering teams experience
  • 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience
  • 8+ years of leading the definition and development of multi tier web services experience
  • Knowledge of engineering practices and patterns for the full software/hardware/networks development life cycle, including coding standards, code reviews, source control management, build processes, testing, certification, and livesite operations
  • Experience partnering with product or program management teams

PREFERRED QUALIFICATIONS

  • Experience in communicating with users, other technical teams, and senior leadership to collect requirements, describe software product features, technical designs, and product strategy
  • Experience in recruiting, hiring, mentoring/coaching and managing teams of Software Engineers to improve their skills, and make them more effective, product software engineers

Benefits & conditions

AD&D insurance, Parental leave, Health insurance, 401(k) matching, Paid time off, Vision insurance, Dental insurance, Flexible spending account Full-time On call Seattle, WA, The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits. USA, WA, Seattle - 184,900.00 - 250,200.00 USD annually

About the company

A NVIDIA UltraServer isn’t a server: it’s 18 compute nodes, 72 Blackwell GPUs, and an NVLink fabric that behaves as one giant, memory-coherent computer. When any part of it fails, you have to diagnose across racks, orchestrate hardware swaps with data center technicians, rebuild the NVLink fabric, re-run burn-in, and hand healthy capacity back to customers training frontier AI models. Every hour an UltraServer sits idle is some of the most sought-after compute on Earth going to waste.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

3:39 min

Addressing code review surrender and process exploitation

Laura Tacho Laura Tacho · World Congress 2026 Europe

51 sec

Repurposing hardware and operating underwater data centers

Chris Heilmann +1 · LIVE

3:49 min

Container hosting options available on Amazon Web Services

Federico Fregosi · World Congress 2022

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · World Congress 2021

56 sec

The hidden costs of delayed peer code reviews

Tim Gilboy Tim Gilboy

Videos

See all

Related articles

See all