AI Accelerator Software Principal Engineer - NPU Full-Stack Integration

Ampere Computing
Portland, OR, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$182,000.0 - $273,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Artificial Neural Networks C++ (Programming Language) Cloud Engineering Profiling Computer Engineering Linux Programming Tools High-Level Architecture Python (Programming Language) Linux Kernel Machine Learning
+9 more
Performance Tuning Software Engineering System Programming Graphics Processing Unit (GPU) Pytorch Deep Learning Information Technology Low Latency Integration Frameworks

Job description

As an AI Accelerator Software Principal Engineer - NPU Full-Stack Integration, you will lead the design and delivery of high-performance, low-latency deep learning inference solutions on the Arm® Ethos -U85. You’ll help advance Ampere’s AI software stack by enabling models with performance and efficiency requirements

You will operate at the intersection of software engineering, performance engineering, and hardware-aware optimization, contributing to the full stack from model execution to accelerator-ready kernel performance.

What You’ll Achieve: * End-to-end deep learning performance acceleration Go deep into the full software/hardware execution stack, including: * + framework integration layers

  • compiler and graph/runtime support
  • runtime libraries and user-mode execution paths
  • compute kernel development
  • profiling, benchmarking, and performance tuning * Model enablement with quality and speed Improve both performance and accuracy for models using popular frameworks, helping deliver production-ready inference behavior in edge devices.

  • Hardware/software co-design and optimization Partner with hardware and platform teams to co-optimize AI execution for better outcomes: *

  • increased throughput
  • reduced latency
  • improved scalability
  • better resource utilization (compute/memory/IO)
  • higher sustained performance under realistic workloads
  • Build state-of-the-art AI software components Contribute to the development of software and hardware AI co-processors/accelerators, delivering reusable libraries, optimized execution paths, and robust integration with existing tooling.

  • Cross-functional collaboration Work closely with cross-functional teams (compiler/runtime, kernels, platform, and product engineering) to integrate AI capabilities into Ampere’s cloud-native processor platforms and accelerators.

Requirements

  • Education and Experience: BS Computer Science, Computer Engineering, Electrical Engineering, or Software Engineering or related technical field & 8 years of related experience; or MS degree & 6 years; or PhD & 3 years. * Understands AOT (Ahead-Of-Time) compilation path in popular frameworks like PyTorch and deployment path like execuTorch in edge environment

  • Linux + accelerator/runtime expertise (preferred): Experience with developing user-mode drivers, runtime libraries, or low-level integration for GPUs or deep learning accelerators in Linux is a plus.
  • Strong systems programming & performance skills: *

  • Expert in Python and C/C++
  • Strong background in performance profiling and tuning (latency/throughput, memory behavior, kernel efficiency)
  • Deep ML understanding: Solid understanding of AI/ML concepts including neural networks and data processing frameworks. Experience with modern deep model architectures such as Transformers and Diffusion models is preferred.
  • Modern AI tooling fluency (preferred): Fluent with modern AI programming tools such as Codex or Claude Code, and comfortable accelerating development workflows.

Benefits & conditions

Pulled from the full job description

  • Flextime
  • Health insurance
  • Vision insurance
  • Dental insurance
  • Paid holidays, At Ampere we believe in taking care of our employees and providing a competitive total rewards package that includes base pay, cash long-term incentive, and comprehensive benefits. The full base pay range for this role is between $182,000 and $273,000, except in the San Francisco Bay Area where the range is between $195,000 and $292,000.

Our benefits include health, wellness, and financial programs that support employees through every stage of life.

Benefit highlights include:

  • Premium medical insurance, dental insurance, vision insurance, as well as income protection and a 401K retirement plan, so that you can feel secure in your health and financial future.
  • Unlimited Flextime and 10+ paid holidays so that you can embrace a healthy work-life balance.
  • A variety of healthy snacks, energizing espresso, and refreshing drinks to keep you fueled and focused throughout the day.

And there is much more than compensation and benefits. At Ampere, we foster an inclusive culture that empowers our employees to do more and grow more. We are excited to share more about our career opportunities with you through the interview process. Our benefits include health, wellness, and financial programs that support employees through every stage of life.

LI-Hybrid

LI-Hybrid#LI-DR

#LI-Hybrid

About the company

Ampere is a semiconductor design company for a new era, leading the future of computing with an innovative approach to CPU design focused on high-performance, energy efficient AI compute.

As a pioneer in the new frontier of energy efficient high-performance computing, Ampere is part of the Softbank Group of companies driving sustainable computing for AI, Cloud, and edge applications.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:20 min

Utilizing AI and hardware acceleration for application code optimization

Stephan Gillich Stephan Gillich · WWC 2024

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · WWC Europe 2026

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · WWC 2025

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

Videos

See all

Related articles

See all