Software Development Engineer I - AI/ML Network Infrastructure, Annapurna Labs

Amazon.com, Inc.
Cupertino, United States of America
2 days ago

Role details

Contract type
Internship / Graduate position
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Junior
Compensation
$ 185K

Job location

Cupertino, United States of America

Tech stack

Artificial Intelligence
Amazon Web Services (AWS)
Amazon Web Services (AWS)
C++
Computer Clusters
Profiling
Protocol Stack
Nvidia CUDA
Computer Engineering
Continuous Integration
Distributed Systems
Memory Management
Fault Tolerance
Python
Linux kernel
Linux System Administration
Message Passing Interface
Network Architecture
Network Programming
Network Service
Open Source Technology
Remote Direct Memory Access
Software Engineering
Toolchain
Multithreading
Grafana
Gpu Programming
Linux Development
Information Technology
Machine Learning Operations

Job description

We're looking for a talented early-career engineer to join our team that owns the network stack for EC2 distributed AI/ML systems. You'll work on software that enables the world's largest AI models to train across massive GPU clusters, developing support for communication libraries and frameworks like NCCL, NVSHMEM, and NIXL., Write high-performance C/C++ code for network communication libraries running on custom AWS hardware

  • Build and maintain infrastructure that monitors functionality and performance of large-scale AI/ML workloads
  • Develop automation using Python and AWS tools (CI/CD, Grafana, Athena) to test, benchmark, and deliver software to customers
  • Design mechanisms to detect functional and performance regressions before they reach production
  • Work across many instance types, software stacks, and Linux environments

Requirements

Bachelor's or Master's degree in Computer Science, Computer Engineering, or related field (recent graduates welcome)

  • Strong proficiency in C/C++
  • Solid coursework or project experience in: 1/ Operating Systems (Linux internals, kernel concepts, memory management) 2/ Parallel Computer Architecture (multi-threading, SIMD, GPU programming, cache coherence) 3/ Distributed Systems (consensus, message passing, fault tolerance, scalability)
  • Familiarity with Linux development environments and toolchains

Preferred Qualifications

  • Internship experience in ML communications, HPC networking, or RDMA/high-speed interconnects
  • Exposure to network programming (sockets, MPI, collective communication patterns)
  • Experience with performance profiling and optimization
  • Familiarity with GPU programming (CUDA) or hardware-software co-design
  • Contributions to open-source projects in systems, networking, or HPC

Benefits & conditions

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, CA, Cupertino - 127,100.00 - 185,000.00 USD annually

Apply for this position