> Markdown version of [/jobs/ext/1182527-ai-accelerator-software-principal-engineer-runtime-library](https://www.wearedevelopers.com/jobs/ext/1182527-ai-accelerator-software-principal-engineer-runtime-library). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Accelerator Software Principal Engineer - Runtime Library - **Company:** Ampere Computing - **Location:** Santa Clara, CA, United States - **Salary:** $182,000.0 - $273,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, C++ (Programming Language), Compilers, Profiling, Computer Programming, Computer Engineering, Extract Transform Load (ETL), Linux, Memory Management, Interoperability, Real-Time Operating Systems, Software Engineering, JavaScript Pagination Plugin, Graphics Processing Unit (GPU), Pytorch, Deep Learning, Information Technology, Low Latency, ONNX (Open Neural Network Exchange) Format - **Published:** July 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=c4e85333f56ba72d ## About the Role * BS Computer Science, Computer Engineering, Electrical Engineering, or Software Engineering or related technical field & 8 years of related experience; or MS degree & 6 years; or PhD & 3 years * Proven experience developing user-mode drivers and/or runtime libraries for GPUs or deep learning accelerators in Linux or RTOS environments. * Strong expertise in C/C++ and systems-level programming (memory, threading, synchronization, performance profiling). * Demonstrated background in AI framework enablement, with hands-on experience in one or more of: * + PyTorch (operator/runtime integration, graph execution, correctness/performance work) + llama.cpp (inference/runtime execution patterns) + ONNX (graph handling, interoperability, execution engines) * Strong performance engineering skills, including profiling/diagnostics and optimization of execution pipelines, data movement, and compute kernels. * Ability to operate effectively in a collaborative environment-owning complex components while partnering with compilers, hardware, and platform teams. ## Description As an AI Accelerator Principal Software Engineer - Runtime Library, you will lead the design, development, and optimization of AI runtime software that enables multiple state-of-the-art deep learning models to run efficiently on Ampere's deep learning accelerators. You will work at the intersection of systems software, performance engineering, and AI enablement, helping deliver high-throughput, low-latency inference and a strong foundation for future model and framework support. What You'll Achieve: * Build and evolve an AI Runtime Library for Ampere accelerators that supports execution, scheduling, and lifecycle management of deep learning workloads across multiple model types and popular frameworks. * Own end-to-end acceleration paths, going deep into the full SW/HW stack-including: * + Inference serving and integration layers + Compiler/runtime interfaces and graph/IR execution flows + Runtime library architecture (APIs, memory management, operators, execution engines) + Communication mechanisms and device/host orchestration * Drive HW/SW co-design and optimization to improve: * + Throughput (tokens/requests per second) + Latency (kernel execution and scheduling efficiency) + Memory efficiency (buffering, paging, reuse, caching) + Overall compute utilization and scaling behavior * Contribute to AI co-processor/accelerator software enablement, partnering closely with hardware and systems teams to ensure runtime and kernel strategies match accelerator capabilities and constraints. * Collaborate cross-functionally to integrate runtime components into Ampere platform stacks, ensuring robust deployment on target environments and consistent performance in production-like workloads. ## Related Videos - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [From Model to Metal: An Open Source Stack for Accelerating Intelligence](https://www.wearedevelopers.com/videos/1636-from-model-to-metal-an-open-source-stack-for-accelerating-intelligence) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Just-in-time Compilation in JVM](https://www.wearedevelopers.com/videos/240-just-in-time-compilation-in-jvm) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)