ML Software Engineer, Data Plane

Amazon.com, Inc.
Cupertino, United States of America
yesterday

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Intermediate
Compensation
$ 224K

Job location

Cupertino, United States of America

Tech stack

Board Bringup
Artificial Intelligence
C++
Code Review
ETL
Distributed Systems
Memory Management
Machine Learning
Open Source Technology
Remote Direct Memory Access
TensorFlow
Software Engineering
Network Switches
Graphics Processing Unit (GPU)
PyTorch
Large Language Models
Model Validation
Parallel Computation
Optimization Algorithms
Build Process
TensorRT
Software Coding
Software Version Control

Job description

The MLIL DataPlane team is looking for a Software Development Engineer to own the design and implementation of our inference data plane. We build the software that makes large models run efficiently on custom hardware - spanning model execution, memory management, data movement, and serving integration. Our work covers the full inference path: integrating serving engines with custom hardware, developing high-performance compute kernels, enabling efficient data movement, and driving models from early validation through production. We operate at frontier scale with large distributed models. This is a ground-up effort with rapidly evolving hardware and software. We are looking for an individual contributor who can write and optimize low-level code for custom hardware, validate model architectures end-to-end, build test and profiling infrastructure, and drive performance across the stack.

Key job responsibilities

  • Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
  • Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware.
  • Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
  • Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
  • Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bring-up.
  • Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
  • Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.

Requirements

Bachelor's degree or equivalent

  • 4+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Knowledge of computer architecture, operating systems, and parallel computing
  • Knowledge of Linux fundamentals
  • Strong proficiency in C/C++
  • Experience developing compute kernels for GPUs, DSPs, or custom accelerators
  • Proven track record of owning and delivering complex software features end-to-end, Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques
  • Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT
  • Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware
  • Experience with distributed systems - collective communication, RDMA, or high-speed interconnect programming
  • Experience with hardware simulation environments and model validation workflows
  • Demonstrated early adopter of AI-assisted development tools - uses LLMs or code-generation agents as part of daily workflow

Benefits & conditions

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, CA, Cupertino - 165,200.00 - 223,600.00 USD annually

Apply for this position