Sr. Software Engineer- AI/ML, AWS Neuron Apps

Amazon.com, Inc.
Seattle, WA, United States
20 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$168,100.0 - $227,400.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services C++ (Programming Language) Code Review Nvidia CUDA Computer Programming Software Debugging Python (Programming Language) Machine Learning Tensorflow Software Engineering Pytorch
+9 more
Large Language Models Parallel Computation Information Technology Build Process Hardware Acceleration Machine Learning Operations Software Coding GPT Software Version Control

Job description

Join the elite team behind AWS Neuron-the software stack powering AWS’s next-generation AI accelerators Inferentia and Trainium. As a Senior Software Engineer in our Machine Learning Applications team, you’ll be at the forefront of deploying and optimizing some of the world’s most sophisticated AI models at unprecedented scale.

What You’ll Impact:

  • Pioneer distributed inference solutions for industry-leading LLMs such as GPT, Llama, Qwen

  • Optimize breakthrough language and vision generative AI models

  • Collaborate directly with silicon architects and compiler teams to push the boundaries of AI acceleration

  • Drive performance benchmarking and tuning that directly impacts millions of inference calls globally

Key job responsibilities

You will drive the Evolution of Distributed AI at AWS Neuron

As a Technical Leader at the forefront of AWS’s AI Accelerator, you’ll architect the bridge between ML frameworks including PyTorch, JAX and AI hardware. This isn’t just about just optimization-it’s about revolutionizing how AI models run at scale.

Technical Impact You’ll Drive:

  • Spearhead distributed inference architecture for PyTorch and JAX using XLA

  • Engineer breakthrough performance optimizations for AWS Trainium and Inferentia

  • Develop ML tools to enhance LLM accuracy and efficiency

Requirements

  • 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • 5+ years of programming experience using Python or C++ and PyTorch.
  • Experience with AI acceleration via quantization, parallelism, model compression, batching, KV caching, vllm serving
  • Experience with accuracy debugging & tooling, performance benchmarking of AI accelerators
  • Fundamentals of Machine learning and deep learning models, their architecture, training and inference lifecycles along with work experience on optimizations for improving the model execution., * Master’s degree in computer science or equivalent
  • Master’s degree in machine learning or equivalent
  • Experience with accuracy debugging & tooling, performance benchmarking of AI accelerators
  • Experience in developing CUDA kernels, HPC and inference optimization, tensors operations

Benefits & conditions

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .

USA, WA, Seattle - 168,100.00 - 227,400.00 USD annually

About the company

Our team puts a high value on work-life balance. It isn’t about how many hours you spend at home or at work; it’s about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives.

Mentorship & Career Growth

Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we’re building an environment that celebrates knowledge sharing and mentorship. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded professional and enable them to take on more complex tasks in the future.

About the team

At AWS Neuron, we’re revolutionizing how the world’s most sophisticated AI models run at scale through Amazon’s next-generation AI accelerators. Operating at the unique intersection of ML frameworks and custom silicon, our team drives innovation from silicon architecture to production software deployment.

We pioneer distributed inference solutions for PyTorch and JAX using XLA, optimize industry-leading LLMs like GPT and Llama, and collaborate directly with silicon architects to influence the future of AI hardware. Our systems handle millions of inference calls daily, while our optimizations directly impact thousands of AWS customers running critical AI workloads.

We’re focused on pushing the boundaries of large language model optimization, distributed inference architecture, and hardware-specific performance tuning. Our deep technical experts transform complex ML challenges into elegant, scalable solutions that define how AI workloads run in production.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:36 min

Exploring high-level Python frameworks for accelerated enterprise artificial intelligence

Paul Graham Paul Graham · LIVE

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

1:39 min

Fundamentals of tensors and the TensorFlow library

Håkan Silfvernagel · LIVE

40 sec

Generative pre-trained transformer models powering code completions

lgonta lgonta +1 · WWC 2024

4:41 min

Replacing PyTorch with ONNX runtime for AWS Lambda deployments

Marek Suppa · LIVE

3:53 min

Architecting machine learning projects with the PAI platform

Qiyang Duan · LIVE

Videos

See all

Related articles

See all