Principal Software Engineer, AI Compute Infrastructure

ARM
Seattle, WA, United States
1 day ago
Apply on www.jofdav.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$262,700.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Cloud Computing Nvidia CUDA Computer Programming Linux Distributed Systems Python (Programming Language) Machine Learning Octopus Deploy Performance Tuning Prometheus Systems Integration
+6 more
Pytorch Large Language Models Grafana Kubernetes TensorRT Terraform

Job description

As a Principal Engineer on the AI Compute Infra team, you will design, build, and operate large-scale infrastructure for AI training, fine-tuning, evaluation, and inference. You will guide work across Kubernetes clusters, accelerator enablement, workload scheduling, high-performance networking, storage, and capacity management, partnering with AI researchers and engineers to improve reliability, performance, scalability, and developer productivity., * Build and operate Kubernetes clusters while improving workload scheduling, topology-aware placement, capacity use, and recovery.

  • Enable new CPU and GPU systems by integrating and validating drivers, networking, storage, monitoring, and health checks.
  • Investigate performance and reliability issues across applications, cloud infrastructure, clusters, and hardware, then turn findings into lasting improvements.
  • Partner with AI teams to understand their workloads and automate cluster provisioning, upgrades, monitoring, and maintenance around their needs.

Necessary Skills and Experience, You will join a driven group committed to developing world-class AI compute infrastructure. We provide a cooperative setting where your ideas can come to life. Your efforts will directly impact the success of our AI projects, guaranteeing smooth operations and outstanding results. Join us and help build the future of AI compute infrastructure!

Requirements

  • 8+ years of experience building or operating cloud, compute, HPC, or distributed infrastructure in a production environment.
  • Programming experience in Go, Python, or another systems language, with an interest in developing reliable infrastructure software.
  • Practical knowledge of Kubernetes, containers, Linux, networking, and storage.
  • Experience supporting GPU, accelerator, or distributed machine-learning workloads.
  • An ability to troubleshoot complex systems and communicate clearly with engineers from different technical backgrounds.

“Preferred” Skills and Experience:

  • Familiarity with Kubernetes scheduling, operators, quotas, or resource management.
  • Experience with NVIDIA technologies such as CUDA, NVLink, NVSwitch, NCCL, EFA, or DCGM.
  • Knowledge of AWS EKS, Terraform, Argo CD, Helm, Prometheus, or Grafana.
  • Familiarity with frameworks such as PyTorch, Ray, vLLM, SGLang, or TensorRT-LLM, or experience qualifying accelerators and tuning distributed workloads.

Benefits & conditions

Salary Range:$262,700-$355,400 per year We value people as individuals and our dedication is to reward people competitively and equitably for the work they do and the skills and experience they bring to Arm. Salary is only one component of Arm’s offering. The total reward package will be shared with candidates during the recruitment and selection process.

Accommodations at Arm

At Arm, we want to build extraordinary teams. If you need an adjustment or an accommodation during the recruitment process, please email accommodations@arm.com. To note, by sending us the requested information, you consent to its use by Arm to arrange for appropriate accommodations. All accommodation or adjustment requests will be treated with confidentiality, and information concerning these requests will only be disclosed as necessary to provide the accommodation. Although this is not an exhaustive list, examples of support include breaks between interviews, having documents read aloud, or office accessibility. Please email us about anything we can do to accommodate you during the recruitment process.

Hybrid Working at Arm

Arm’s approach to hybrid working is designed to create a working environment that supports both high performance and personal wellbeing. We believe in bringing people together face to face to enable us to work at pace, whilst recognizing the value of flexibility. Within that framework, we empower groups/teams to determine their own hybrid working patterns, depending on the work and the team’s needs. Details of what this means for each role will be shared upon application. In some cases, the flexibility we can offer is limited by local legal, regulatory, tax, or other considerations, and where this is the case, we will collaborate with you to find the best solution. Please talk to us to find out more about what this could look like for you.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jofdav.com
Prepare application

Inside Arm

Culture, engineering, and team stories

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · World Congress 2025

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all