> Markdown version of [/jobs/ext/1215790-ai-accelerator-software-principal-engineer-npu-full-stack-integration](https://www.wearedevelopers.com/jobs/ext/1215790-ai-accelerator-software-principal-engineer-npu-full-stack-integration). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Accelerator Software Principal Engineer - NPU Full-Stack Integration - **Company:** Ampere Computing - **Location:** Santa Clara, CA, United States - **Salary:** $182,000.0 - $273,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Artificial Neural Networks, C++ (Programming Language), Cloud Engineering, Profiling, Computer Engineering, Linux, Programming Tools, High-Level Architecture, Python (Programming Language), Linux Kernel, Machine Learning, Performance Tuning, Software Engineering, System Programming, Graphics Processing Unit (GPU), Pytorch, Deep Learning, Information Technology, Low Latency, Integration Frameworks - **Published:** July 9, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=90868970d90f700a ## About the Role * Education and Experience: BS Computer Science, Computer Engineering, Electrical Engineering, or Software Engineering or related technical field & 8 years of related experience; or MS degree & 6 years; or PhD & 3 years. * Understands AOT (Ahead-Of-Time) compilation path in popular frameworks like PyTorch and deployment path like execuTorch in edge environment * Linux + accelerator/runtime expertise (preferred): Experience with developing user-mode drivers, runtime libraries, or low-level integration for GPUs or deep learning accelerators in Linux is a plus. * Strong systems programming & performance skills: * + Expert in Python and C/C++ + Strong background in performance profiling and tuning (latency/throughput, memory behavior, kernel efficiency) * Deep ML understanding: Solid understanding of AI/ML concepts including neural networks and data processing frameworks. Experience with modern deep model architectures such as Transformers and Diffusion models is preferred. * Modern AI tooling fluency (preferred): Fluent with modern AI programming tools such as Codex or Claude Code, and comfortable accelerating development workflows. ## Description As an AI Accelerator Software Principal Engineer - NPU Full-Stack Integration, you will lead the design and delivery of high-performance, low-latency deep learning inference solutions on the ArmĀ® Ethos -U85. You'll help advance Ampere's AI software stack by enabling models with performance and efficiency requirements You will operate at the intersection of software engineering, performance engineering, and hardware-aware optimization, contributing to the full stack from model execution to accelerator-ready kernel performance. What You'll Achieve: * End-to-end deep learning performance acceleration Go deep into the full software/hardware execution stack, including: * + framework integration layers + compiler and graph/runtime support + runtime libraries and user-mode execution paths + compute kernel development + profiling, benchmarking, and performance tuning * Model enablement with quality and speed Improve both performance and accuracy for models using popular frameworks, helping deliver production-ready inference behavior in edge devices. * Hardware/software co-design and optimization Partner with hardware and platform teams to co-optimize AI execution for better outcomes: * + increased throughput + reduced latency + improved scalability + better resource utilization (compute/memory/IO) + higher sustained performance under realistic workloads * Build state-of-the-art AI software components Contribute to the development of software and hardware AI co-processors/accelerators, delivering reusable libraries, optimized execution paths, and robust integration with existing tooling. * Cross-functional collaboration Work closely with cross-functional teams (compiler/runtime, kernels, platform, and product engineering) to integrate AI capabilities into Ampere's cloud-native processor platforms and accelerators. ## Related Videos - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [From Model to Metal: An Open Source Stack for Accelerating Intelligence](https://www.wearedevelopers.com/videos/1636-from-model-to-metal-an-open-source-stack-for-accelerating-intelligence) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [WWC24 - Ankit Patel - Unlocking the Future Breakthrough Application Performance and Capabilities with NVIDIA](https://www.wearedevelopers.com/videos/920-wwc24-ankit-patel-unlocking-the-future-breakthrough-application-performance-and-capabilities-with-nvidia) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care)