Software Development Engineer I - AI/ML Network Infrastructure, Annapurna Labs

Amazon.com, Inc.
Cupertino, CA, United States
21 days ago

Role details

Contract type
Internship / Graduate position
Employment type
Full-time (> 32 hours)
Experience level
Starter
Compensation
$127,100.0 - $185,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud C++ (Programming Language) Computer Clusters Profiling Protocol Stack Nvidia CUDA Computer Engineering Continuous Integration Distributed Systems Memory Management
+18 more
Fault Tolerance Python (Programming Language) Linux Kernel Linux System Administration Message Passing Interface Network Architecture Network Programming Network Service Open Source Technology Remote Direct Memory Access Software Engineering Toolchain Multithreading Grafana Gpu Programming Linux Development Information Technology Machine Learning Operations

Job description

We’re looking for a talented early-career engineer to join our team that owns the network stack for EC2 distributed AI/ML systems. You’ll work on software that enables the world’s largest AI models to train across massive GPU clusters, developing support for communication libraries and frameworks like NCCL, NVSHMEM, and NIXL., Write high-performance C/C++ code for network communication libraries running on custom AWS hardware

  • Build and maintain infrastructure that monitors functionality and performance of large-scale AI/ML workloads
  • Develop automation using Python and AWS tools (CI/CD, Grafana, Athena) to test, benchmark, and deliver software to customers
  • Design mechanisms to detect functional and performance regressions before they reach production
  • Work across many instance types, software stacks, and Linux environments

Requirements

Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or related field (recent graduates welcome)

  • Strong proficiency in C/C++
  • Solid coursework or project experience in: 1/ Operating Systems (Linux internals, kernel concepts, memory management) 2/ Parallel Computer Architecture (multi-threading, SIMD, GPU programming, cache coherence) 3/ Distributed Systems (consensus, message passing, fault tolerance, scalability)
  • Familiarity with Linux development environments and toolchains

Preferred Qualifications

  • Internship experience in ML communications, HPC networking, or RDMA/high-speed interconnects
  • Exposure to network programming (sockets, MPI, collective communication patterns)
  • Experience with performance profiling and optimization
  • Familiarity with GPU programming (CUDA) or hardware-software co-design
  • Contributions to open-source projects in systems, networking, or HPC

Benefits & conditions

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, CA, Cupertino - 127,100.00 - 185,000.00 USD annually

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.amazon.jobs

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:12 min

Integrating toolchains to improve engineering efficiency and culture

Amir Friedman Amir Friedman +2 · WWC 2025

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · WWC Europe 2026

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · WWC 2025

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

Videos

See all

Related articles

See all