Software Development Engineer II, Annapurna Labs

Amazon.com, Inc.
Austin, TX, United States
4 days ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Internship / Graduate position
Employment type
Full-time (> 32 hours)
Experience required
1 year minimum
Compensation
$143,700.0 - $194,400.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Artificial Intelligence Artificial Neural Networks BIOS C++ (Programming Language) Cloud Computing Profiling Code Review Computer Programming Continuous Integration Data Visualization Software Debugging
+22 more
Software Design Patterns Dynamic Random-Access Memory Emulators Firmware Hardware Interface Design Python (Programming Language) Lua (Scripting Language) Machine Learning Regression Analysis PCI Express Software Engineering SystemVerilog Rust (Programming Language) Graphics Processing Unit (GPU) Pytorch Caching Information Technology AWS Data Analytics Build Process Software Coding Software Version Control Golang

Job description

As a Validation Engineer on our Machine Learning Acceleration team, you’ll own critical validation aspects across the entire product development lifecycle-from early design validation through emulation, silicon bring-up, post-silicon validation, and ongoing support of production systems deployed in AWS data centers. You’ll collaborate deeply with architecture, RTL design, design verification, firmware, and software teams to ensure our next-generation AI/ML accelerators meet the highest standards of quality and performance. This role requires bridging multiple domains-from low-level hardware interfaces to high-level ML workloads-to deliver exceptional results., * Developing comprehensive validation strategies and detailed test plans covering functional, performance, power, and stress testing from silicon bring-up to product release

  • Executing complex test plans from RTL simulation and emulation environments through physical silicon validation
  • Conducting hands-on silicon bring-up and debug in the lab using oscilloscopes, logic analyzers, and protocol analyzers
  • Validating ML accelerator performance, accuracy, and reliability using real-world neural network workloads
  • Building test infrastructure, CI/CD, and automated regression frameworks to enable efficient validation at scale
  • Collaborating across architecture, design, firmware, and software teams to triage failures and drive root cause analysis to closure
  • Reviewing test results, identifying patterns, and providing feedback to improve design quality and validation coverage
  • Supporting production systems in AWS data centers and addressing field issues as they arise

Requirements

  • Strong programming skills (Python, Lua, C/C++, Rust, Go, etc)
  • A solid understanding of computer architecture
  • Experience with AWS services, cloud infrastructure, firmware development (BIOS, BMC, drivers)
  • Validation experience in any of these areas: PCIe, HBM, GPUs, neural networks, ML HW architecture, and/or CI/CD
  • Familiarity with the validation lifecycle from RTL simulation (SystemVerilog/UVM, VCS, Questa, Xcelium) and emulation (Palladium, Zebu, Veloce) through silicon failure analysis and debug, * Experience programming with at least one software programming language
  • 3+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • Experience with Machine Learning and Large Language Model fundamentals, including architecture, training/inference lifecycles, and optimization of model execution, or experience working with PyTorch or JAX software
  • Bachelor’s degree in computer science, engineering, mathematics or equivalent, or experience in Java, C++, Python, or a related language
  • Experience in debugging, profiling, and implementing software engineering best practices in large-scale systems, or experience in development in the last 3 years
  • Strong understanding of computer architecture fundamentals including memory hierarchies (caches, DRAM, HBM), compute pipelines, and interconnect topologies
  • Experience applying statistical methods, regression analysis, and data visualization techniques to interpret performance data and drive optimization decisions, * 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Bachelor’s degree in computer science or equivalent
  • Experience with machine learning (ML) tools and methods
  • 1+ years of continuous integration and continuous delivery (CI/CD), or 1+ years of continuous integration and continuous delivery (CI/CD) experience
  • Experience with EDA Simulations or Emulation

Benefits & conditions

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .

USA, TX, Austin - 143,700.00 - 194,400.00 USD annually

About the company

Annapurna Labs, an AWS organization with development centers in the U.S. and Israel, builds custom silicon and software for AWS customers. Our team combines cloud-scale innovation with world-class expertise across silicon engineering, hardware design, verification, software, and operations to tackle technical challenges that have never been seen before.

Join our Silicon Validation team to validate next-generation machine learning accelerators that power AWS’s cloud computing infrastructure. You’ll work in a fast-paced, startup-like environment alongside some of the brightest minds in the industry on cutting-edge, internet-scale technology that directly impacts how customers use Machine Learning acceleration. We are changing the landscape of cloud infrastructure by accelerating the development of custom silicon by moving beyond traditional partnerships to dominate in AI training and inference

Your work will span validation of the complete vertical stack-silicon, PCB, high-speed components (HBM, PCIe, chip-to-chip), inter-system connections, and system-to-system interfaces. You’ll dive deep into new technology hardware components and scaling technologies that power our Machine Learning boards and servers at scale, ensuring every component of our hardware and software comes together into products our customers rely on.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · World Congress 2026 Europe

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

41 sec

Massive client data loss and bio-digital storage

Chris Heilmann Chris Heilmann +1 · LIVE

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

2:33 min

Maintaining prompt structures for prefix caching

Douglas Reiser Douglas Reiser · Europe 2026 Virtual

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

Videos

See all

Related articles

See all