> Markdown version of [/jobs/ext/2847708-staff-software-engineer-ai-inference-runtime](https://www.wearedevelopers.com/jobs/ext/2847708-staff-software-engineer-ai-inference-runtime). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer, AI Inference Runtime - **Company:** ARM - **Location:** Seattle, WA, United States - **Experience:** Expert - **Salary:** $209,100.0 - $282,900.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Compilers, Computer Programming, Extract Transform Load (ETL), Software Debugging, Programming Tools, Memory Management, Python (Programming Language), Linux Kernel, Open Source Technology, AI Infrastructure, Rust (Programming Language), Concurrency, Parallel Computation, AI Platforms, Low Latency, Machine Learning Operations - **Published:** September 11, 2026 - **Apply:** https://www.jofdav.com/jobs/59633961-staff-software-engineer-ai-inference-runtime ## About the Role * 5+ years of experience, or equivalent demonstrated impact, in ML systems, high-performance systems, compilers, kernel development, or production AI inference. * Deep understanding of modern AI inference, including model execution, Attention, MoE, batching, prioritisation, and KV-cache behavior. * Strong programming skills in C++, Rust, Python, or a comparable language, with knowledge of concurrency, parallel programming, handling of memory resources, and data movement. * Proven ability to profile, debug, and optimize performance across kernels, runtimes, frameworks, operating systems, and hardware. "Nice To Have" Skills and Experience : * Experience developing or modifying inference schedulers, cache managers, batching systems, disaggregated or distributed execution paths. * Experience optimizing kernels using accelerator programming tools, assembly, or intrinsics, including attention, matrix multiplication, operator fusion, and low-precision execution. * Familiarity with model parallelism, collective communication, high-performance networking, compilers, or graph optimization. * Contributions to open-source ML runtimes, frameworks, compilers, or kernel libraries. ## Description As a Software Engineer on our AI Inference Runtime team, you will set technical direction for critical components of distributed Inference runtime for running SOTA AI Models. You will lead hands-on work across scheduling, batching, KV-cache management, memory allocation, distributed workload execution, kernel development and optimization, and performance benchmarking and analysis. Your work will directly influence how efficiently new models use available compute. Partnering with our AI Infrastructure, compute, and product teams to enhance the performance and efficiency of Arm's AI platform., * Define the architecture, interfaces, and roadmap for AI inference runtime capabilities, including abstractions that support evolving models, workloads, and compute platforms. * Enable new model architectures end to end through operator support, production validation, and optimization of scheduling, batching, model execution, memory management, and KV-cache efficiency. * Profile system bottlenecks and develop optimized kernels and data-movement paths across compute, memory, networking, and framework integration. * Evaluate new inference techniques and build benchmarking, regression, validation, and safe-rollout systems to improve latency, throughput, reliability, and resource efficiency. * Partner with cloud, framework, compiler, hardware, and research teams; lead technical reviews, mentor engineers, and establish meticulous performance-engineering practices. ## About Arm Arm is the industry’s highest-performing and most power-efficient compute platform with unmatched scale that touches 100 percent of the connected global population. To meet the insatiable demand for compute, Arm is delivering advanced solutions that allow the world’s leading technology companies to unleash the unprecedented experiences and capabilities of AI. Together with the world’s largest computing ecosystem and 22 million software developers, we are building the future of AI on Arm. [Company profile](https://www.wearedevelopers.com/companies/3413-arm) ### More Jobs at Arm - [Staff Software Engineer, AI Compute Infrastructure](https://www.wearedevelopers.com/jobs/ext/2847711-staff-software-engineer-ai-compute-infrastructure) - [Staff Memory Controller Performance Architect](https://www.wearedevelopers.com/jobs/ext/2844510-staff-memory-controller-performance-architect) - [Staff Software Engineer, AI Inference Cloud](https://www.wearedevelopers.com/jobs/ext/2844507-staff-software-engineer-ai-inference-cloud) - [Principal Software Engineer, AI Compute Platform](https://www.wearedevelopers.com/jobs/ext/2847710-principal-software-engineer-ai-compute-platform) - [Principal Software Engineer, AI Compute Infrastructure](https://www.wearedevelopers.com/jobs/ext/2847709-principal-software-engineer-ai-compute-infrastructure) ## Related Videos - [Unleashing the Full Potential of the Arm Architecture – Write Once, Deploy Anywhere](https://www.wearedevelopers.com/videos/940-unleashing-the-full-potential-of-the-arm-architecture-write-once-deploy-anywhere) - [Just-in-time Compilation in JVM](https://www.wearedevelopers.com/videos/240-just-in-time-compilation-in-jvm) - [Swapping Low Latency Data Storage Under High Load](https://www.wearedevelopers.com/videos/746-swapping-low-latency-data-storage-under-high-load) - [Concurrency with Go](https://www.wearedevelopers.com/videos/191-concurrency-with-go) - [C++ in constrained environments](https://www.wearedevelopers.com/videos/441-c-in-constrained-environments) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)