Machine Learning Engineer, AWS Neuron Inference, Annapurna ML

Amazon.com, Inc.
Seattle, WA, United States
about 1 month ago

Role details

Contract type
Internship / Graduate position
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$104,000.0 - $145,600.0
Working hours
Regular working hours

Tech stack

Adobe InDesign Artificial Intelligence Amazon Web Services Code Review Computer Programming Software Design Patterns Integrated Development Environments Python (Programming Language) Machine Learning Software Engineering Pytorch Large Language Models
+5 more
Information Technology Build Process Software Coding GPT Software Version Control

Job description

AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machine learning accelerators and the Trn2 and future Trn3 servers that use them. This role is for a software engineer in the Machine Learning Applications (ML Apps) team for AWS Neuron. This role develops, enables and performance tunes building blocks for all key ML model families, including Llama3, GPT OSS, Qwen3, DeepSeek and beyond. The Neuron Inference Technology team works side by side with the Inference Model Enablement, compiler runtime engineers to create, build and tune high-performance distributed inference solutions for the latest generation Trainium accelerators. Experience optimizing LLM inference performance with kernels, Python, PyTorch or JAX is a must., This team develops optimized building blocks for the Neuron distributed inference library, tuning them to ensure highest performance and maximize efficiency running on Trn2 and Trn3 servers. A day in the life As you develop technology components, you’ll create metrics, implement automation and other improvements, and resolve the root cause of software defects. You’ll also participate in design discussions, code review, and communicate with internal and external stakeholders. You will work cross-functionally with teams across Neufon in a fast-paced startup-like development environment, where we constantly stay on top of the latest priorities as the AI landscape evolves. About the team Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we’re building an environment that celebrates knowledge-sharing and mentorship. Our senior members enjoy one-on-one mentoring and thorough, but kind, code reviews. We care about your career growth and strive to assign projects that help our team members develop your engineering expertise so you feel empowered to take on more complex tasks in the future.

Requirements

3+ years of non-internship professional software development experience

  • 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • Experience programming with at least one software programming language Preferred Qualifications

  • 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Bachelor’s degree in computer science or equivalent

Benefits & conditions

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at . USA, WA, Seattle - 143,700.00 - 194,400.00 USD annually

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · WWC 2024

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

3:39 min

Addressing code review surrender and process exploitation

Laura Tacho Laura Tacho · WWC Europe 2026

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

3:53 min

Architecting machine learning projects with the PAI platform

Qiyang Duan · LIVE

51 sec

Assessing GPT-4o performance for pull request feedback

Merrill Lutsky Merrill Lutsky · WWC 2025

Videos

See all

Related articles

See all