> Markdown version of [/jobs/ext/2854957-principal-software-engineer-ai-inference-cloud](https://www.wearedevelopers.com/jobs/ext/2854957-principal-software-engineer-ai-inference-cloud). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Software Engineer, AI Inference Cloud - **Company:** ARM - **Location:** Seattle, WA, United States - **Experience:** Expert - **Salary:** $262,700.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, C++ (Programming Language), Cloud Computing, Computer Programming, Software Debugging, Distributed Systems, Python (Programming Language), AI Infrastructure, Rust (Programming Language), Pytorch, Large Language Models, AI Platforms, Kubernetes, Machine Learning Operations, TensorRT - **Published:** September 11, 2026 - **Apply:** https://www.jofdav.com/jobs/59633970-principal-software-engineer-ai-inference-cloud ## About the Role * 8+ years of experience, or equivalent proven impact, building distributed systems, cloud platforms, or production infrastructure. * Deep production experience with Kubernetes, including controllers, operators, scheduling, resource management, networking, and workload lifecycle management. * Strong software and production engineering skills, including hands-on programming in Go, C++, Rust, Python, or a similar language, and experience with reliable services, APIs, concurrency, observability, deployment safety, capacity planning, and incident response. * A track record of leading complex technical initiatives while remaining hands-on, including architecture, implementation, debugging, mentoring, and influencing technical direction across teams. * Ability to troubleshoot complex systems and communicate clearly with engineers from different technical backgrounds. Preferred Skills and Experience: * Experience with AI infrastructure, model serving, or accelerator-backed workloads. * Familiarity with frameworks such as PyTorch, Ray, vLLM, SGLang, or TensorRT-LLM, or experience qualifying accelerators and tuning distributed workloads. * Knowledge of inference performance and resource-efficiency considerations. ## Description As a Principal Engineer on Arm's AI Inference Cloud team, you will shape the technical direction and develop highly available, scalable services for running AI inference workloads. You will guide architecture and actively contribute to development across Kubernetes orchestration, workload management, service delivery, and observability. Partnering with AI compute, Inference Runtime, and product teams to enhance the performance and usability of Arm's AI platform., * Define and build the architecture for cloud-based AI inference services. * Develop Kubernetes controllers and platform capabilities supporting workload deployment, scheduling, recovery, scaling, upgrades, and lifecycle management. * Establish production practices for health validation, progressive rollout, rollback, observability, and service objectives. Improve platform reliability, scalability, performance, and resource efficiency. * Lead production readiness reviews and resolve complex issues across services, Kubernetes, networking, and compute infrastructure. Turn incidents and operational bottlenecks into durable platform improvements. * Lead design and build reviews, mentor engineers, and drive technical alignment across teams. Necessary Skills and Experience ## About Arm Arm is the industry’s highest-performing and most power-efficient compute platform with unmatched scale that touches 100 percent of the connected global population. To meet the insatiable demand for compute, Arm is delivering advanced solutions that allow the world’s leading technology companies to unleash the unprecedented experiences and capabilities of AI. Together with the world’s largest computing ecosystem and 22 million software developers, we are building the future of AI on Arm. [Company profile](https://www.wearedevelopers.com/companies/3413-arm) ### More Jobs at Arm - [Staff Software Engineer, AI Compute Infrastructure](https://www.wearedevelopers.com/jobs/ext/2847711-staff-software-engineer-ai-compute-infrastructure) - [Staff Memory Controller Performance Architect](https://www.wearedevelopers.com/jobs/ext/2844510-staff-memory-controller-performance-architect) - [Staff Software Engineer, AI Inference Cloud](https://www.wearedevelopers.com/jobs/ext/2844507-staff-software-engineer-ai-inference-cloud) - [Principal Software Engineer, AI Compute Platform](https://www.wearedevelopers.com/jobs/ext/2847710-principal-software-engineer-ai-compute-platform) - [Staff Software Engineer, AI Inference Runtime](https://www.wearedevelopers.com/jobs/ext/2847708-staff-software-engineer-ai-inference-runtime) ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) ## Related Articles - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)