> Markdown version of [/jobs/ext/2847711-staff-software-engineer-ai-compute-infrastructure](https://www.wearedevelopers.com/jobs/ext/2847711-staff-software-engineer-ai-compute-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer, AI Compute Infrastructure - **Company:** ARM - **Location:** Seattle, WA, United States - **Experience:** Expert - **Salary:** $209,100.0 - $282,900.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Cloud Computing, Nvidia CUDA, Computer Programming, Linux, Distributed Systems, Python (Programming Language), Machine Learning, Octopus Deploy, Performance Tuning, Prometheus, Systems Integration, Pytorch, Large Language Models, Grafana, Kubernetes, TensorRT, Terraform - **Published:** September 11, 2026 - **Apply:** https://www.jofdav.com/jobs/59633962-staff-software-engineer-ai-compute-infrastructure ## About the Role * 5+ years of experience building or operating cloud, compute, HPC, or distributed infrastructure in a production environment. * Programming experience in Go, Python, or another systems language, with an interest in developing reliable infrastructure software. * Practical knowledge of Kubernetes, containers, Linux, networking, and storage. * Experience supporting GPU, accelerator, or distributed machine-learning workloads. * An ability to troubleshoot complex systems and communicate clearly with engineers from different technical backgrounds. "Preferred" Skills and Experience: * Familiarity with Kubernetes scheduling, operators, quotas, or resource management. * Experience with NVIDIA technologies such as CUDA, NVLink, NVSwitch, NCCL, EFA, or DCGM. * Knowledge of AWS EKS, Terraform, Argo CD, Helm, Prometheus, or Grafana. * Familiarity with frameworks such as PyTorch, Ray, vLLM, SGLang, or TensorRT-LLM, or experience qualifying accelerators and tuning distributed workloads. ## Description As an engineer on the AI Compute Infra team, you will design, build, and operate large-scale infrastructure for AI training, fine-tuning, evaluation, and inference. You will work across Kubernetes clusters, accelerator enablement, workload scheduling, high-performance networking, storage, and capacity management, partnering with AI researchers and engineers to improve reliability, performance, scalability, and developer productivity., * Build and operate Kubernetes clusters while improving workload scheduling, topology-aware placement, capacity use, and recovery. * Enable new CPU and GPU systems by integrating and validating drivers, networking, storage, monitoring, and health checks. * Investigate performance and reliability issues across applications, cloud infrastructure, clusters, and hardware, then turn findings into lasting improvements. * Partner with AI teams to understand their workloads and automate cluster provisioning, upgrades, monitoring, and maintenance around their needs., You will join a driven group committed to developing world-class AI compute infrastructure. We provide a cooperative setting where your ideas can come to life. Your efforts will directly impact the success of our AI projects, guaranteeing smooth operations and outstanding results. Join us and help build the future of AI compute infrastructure! ## About Arm Arm is the industry’s highest-performing and most power-efficient compute platform with unmatched scale that touches 100 percent of the connected global population. To meet the insatiable demand for compute, Arm is delivering advanced solutions that allow the world’s leading technology companies to unleash the unprecedented experiences and capabilities of AI. Together with the world’s largest computing ecosystem and 22 million software developers, we are building the future of AI on Arm. [Company profile](https://www.wearedevelopers.com/companies/3413-arm) ### More Jobs at Arm - [Staff Software Engineer, AI Inference Runtime](https://www.wearedevelopers.com/jobs/ext/2847708-staff-software-engineer-ai-inference-runtime) - [Staff Memory Controller Performance Architect](https://www.wearedevelopers.com/jobs/ext/2844510-staff-memory-controller-performance-architect) - [Staff Software Engineer, AI Inference Cloud](https://www.wearedevelopers.com/jobs/ext/2844507-staff-software-engineer-ai-inference-cloud) - [Principal Software Engineer, AI Compute Platform](https://www.wearedevelopers.com/jobs/ext/2847710-principal-software-engineer-ai-compute-platform) - [Principal Software Engineer, AI Compute Infrastructure](https://www.wearedevelopers.com/jobs/ext/2847709-principal-software-engineer-ai-compute-infrastructure) ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Unleashing the Full Potential of the Arm Architecture – Write Once, Deploy Anywhere](https://www.wearedevelopers.com/videos/940-unleashing-the-full-potential-of-the-arm-architecture-write-once-deploy-anywhere) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)