Senior Software Development Engineer AI/ML, Inference Model Enablement

Annapurna Labs
Cupertino, CA, United States
1 day ago
Apply on www.careerboard.com
Prepare application

Role details

Contract type
Internship / Graduate position
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Java (Programming Language) Artificial Intelligence Amazon Web Services C Sharp (Programming Language) C++ (Programming Language) Code Review Computer Programming Software Debugging Software Design Patterns Machine Learning Object-Oriented Software Development Open Source Technology
+8 more
Performance Tuning Software Engineering Large Language Models Reliability of Systems Information Technology Build Process Software Coding Software Version Control

Job description

Applicants must be eligible to work in the specified location We develop AWS Neuron, the complete software stack for Trainium, Amazon’s custom cloud-scale machine learning accelerators. Join us to optimize the latest models to run really fast on the Trainium hardware.

As a Sr. Software Development Engineer on the Inference Model Enablement team, you will onboard and optimize state-of-the-art open-source and customer LLMs, both dense and MoE, for inference on Trainium accelerators. You will also drive improvements in model enablement speed and experience, while advancing inference usability and quality through inference features, infrastructure optimization, tools, and automation.

Key job responsibilities Deliver high-performance models using distributed inference libraries Drive technical excellence in performance optimization and system reliability across the Neuron ecosystem Mentor team members and provide technical leadership across multiple work streams Drive architectural decisions that impact the entire Neuron serving stack Collaborate with customers, product owners, and engineering teams to define technical strategy Author technical documentation, design proposals, and architectural guidelines

A day in the life You’ll lead critical technical initiatives while mentoring team members. You’ll collaborate with cross-functional teams of applied scientists, system engineers, and product managers to architect and deliver state-of-the-art inference capabilities. Your day might involve:

Leading design reviews and architectural discussions Debugging complex performance issues across the stack in collaboration with the compiler and runtime teams Mentoring junior engineers on system design and model optimization across model enablement teams Driving technical decisions that shape the future of Neuron’s inference stack

Requirements

BASIC QUALIFICATIONS - 5+ years of programming using a modern programming language such as Java, C++, or C#, including object-oriented design experience

  • 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • 5+ years of non-internship professional software development experience
  • Experience as a mentor, tech lead or leading an engineering team PREFERRED QUALIFICATIONS - Master’s degree in computer science or equivalent

  • Experience with Machine Learning and Large Language Model fundamentals, including architecture, training/inference lifecycles, and optimization of model execution

Benefits & conditions

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerboard.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:39 min

Addressing code review surrender and process exploitation

Laura Tacho Laura Tacho · World Congress 2026 Europe

2:59 min

Generating a Swagger JSON file during the build process

Roman Alexis Anastasini · World Congress 2021

2:33 min

Defining hybrid development and no-code software paradigms

Mark Piller · LIVE

1:49 min

Augmenting junior and principal engineering roles with AI

Neel Sundaresan Neel Sundaresan +1 · World Congress 2026 Europe

56 sec

The hidden costs of delayed peer code reviews

Tim Gilboy Tim Gilboy

Videos

See all

Related articles

See all