Lead Software Engineer, ML Network Stack - Annapurna Labs

Amazon.com, Inc.
Cupertino, CA, United States
about 1 month ago
Apply on www.jobmonkeyjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$193,300.0 - $261,500.0
Working hours
Regular working hours

Tech stack

Computer-Aided Design Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud C++ (Programming Language) Code Review Protocol Stack Nvidia CUDA Software Design Patterns Linux Machine Learning Remote Direct Memory Access
+9 more
Software Engineering Graphics Processing Unit (GPU) Session Description Protocol Security Descriptions (SDES) Information Technology Build Process Machine Learning Operations Software Coding Software Version Control Programming Languages

Job description

We are seeking an experienced engineer and technical leader to join our team that owns the network stack for EC2 distributed AI/ML systems. The team develops support for a variety of frameworks and communication libraries including NCCL, NVSHMEM, NIXL, NCCL GIN, and CUDA kernels. Solid knowledge of Linux, networking, and performant coding is important. Experience with embedded systems is valued, and experience with high-speed networking or HPC/RDMA interconnects is highly valued.

If you like solving hard problems, want to work with HPC and ML customers, iterate fast and deliver meaningful solutions at scale, then come join us! This truly is a role at the forefront of AI/ML-you’ll be working on features for the largest clusters, with the largest customers, for the largest AI models.

The organization you would be joining is Annapurna Labs, an integral part of AWS that develops hardware and software components that are critical building blocks for EC2 infrastructure. Every instance in EC2 is running some type of hardware designed by Annapurna Labs. We specialize in designing software, systems, and chips that optimize the AWS customer experience., Be the Leader which works across the NVIDIA ML communication stack for enabling NVIDIA GPUs to work with the AWS EC2 machines. Knowledge of ML Networking, ML applications, and frameworks will be highly regarded. You’ll be leading senior, mid-level, and junior SDEs and directing work to ensure the team delivers functions and features required for the latest and largest ML workloads.

Requirements

5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience

  • 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Experience as a mentor, tech lead or leading an engineering team, or experience in development in the last 3 years
  • 5+ years of experience with programming language: C or C++

Preferred Qualifications

  • Bachelor’s degree in computer science or equivalent
  • Experience working with ML Communication Libraries or RDMA Networking

Benefits & conditions

Come be a part of an organization that has developed some of the foundational technologies in AWS and EC2. Annapurna Labs’ inventions are routinely mentioned in keynotes and all of re:Invent. Do you recognize Nitro, Trainium, EFA, ENA-X, Inferentia, Graviton? Join us and help create the future. We are seeking Senior SDE’s interested in moving to an SDM role after ramp.

About the team Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we’re building an environment that celebrates knowledge-sharing and mentorship. Our senior members enjoy one-on-one mentoring and thorough, but kind, code reviews. We care about your career growth and strive to assign projects that help our team members develop your engineering expertise so you feel empowered to take on more complex tasks in the future.

Diverse Experiences Amazon values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying.

Work/Life Balance We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud.

Inclusive Team Culture Here at AWS, it’s in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences.

Mentorship and Career Growth We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional., The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, CA, Cupertino - 193,300.00 - 261,500.00 USD annually

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jobmonkeyjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · World Congress 2021

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:39 min

Addressing code review surrender and process exploitation

Laura Tacho Laura Tacho · World Congress 2026 Europe

6:21 min

Previewing upcoming hardware acceleration capabilities for Python environments

Chris Heilmann +2 · LIVE

3:30 min

Transitioning from CUDA software architect to user

Stephen Jones · Coffee With Developers

1:37 min

Accelerating compute with focused developer tools

Julia Koch Julia Koch +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all