Principal Software Engineer, AI Compute Platform

ARM
Seattle, WA, United States
1 day ago
Apply on www.jofdav.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$262,700.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Computing Platforms Cloud Computing Computer Programming Relational Databases Programming Tools Distributed Computing Environment Distributed Systems Python (Programming Language) PostgreSQL Machine Learning
+8 more
Prometheus Pytorch Grafana Backend Build Management AI Platforms Kubernetes Apache Kafka

Job description

As a Principal Software Engineer on the AI Compute Platform team, you will design and build a secure, reliable, and easy-to-use platform for running AI workloads at scale. You will develop control-plane services, APIs, orchestration systems, and developer tools for distributed training and inference, working closely with compute infrastructure teams, AI engineers, and internal users to improve reliability, scalability, and developer productivity across Arm’s AI workloads., * Build scalable backend services, APIs, and asynchronous workflows for managing AI workloads.

  • Develop Kubernetes controllers that manage workload submission, recovery, and long-running service lifecycles.
  • Deliver secure multi-tenant capabilities covering access, secrets, quotas, and auditability.
  • Improve platform scalability, observability, reliability, and developer workflows.
  • Partner across teams to deliver distributed training, evaluation, and inference capabilities while strengthening engineering quality.

Requirements

  • 8+ years of experience building distributed systems, cloud platforms, or production backend services
  • Experience with distributed ML training or inference systems
  • Strong programming skills in Go, Python, or similar systems languages
  • Practical knowledge of APIs, Kubernetes, containers, asynchronous processing, and relational databases.
  • Ability to build clear interfaces and simplify infrastructure for engineers and researchers.
  • Clear communication and a collaborative approach to problem-solving.

Desired Skills and Experience:

  • Experience with Kubernetes operators, schedulers, resource management, or GitOps or equivalent experience
  • Knowledge of RPC systems, PostgreSQL, Kafka, OpenTelemetry, Prometheus and Grafana
  • Familiarity with distributed AI platforms, inference serving, or frameworks such as PyTorch, Ray, or vLLM
  • Experience delivering developer tools or platform capabilities from early build through production
  • Technical leadership, mentoring experience, or an advanced degree or equivalent experience in a related field

Benefits & conditions

Salary Range:$262,700-$355,400 per year We value people as individuals and our dedication is to reward people competitively and equitably for the work they do and the skills and experience they bring to Arm. Salary is only one component of Arm’s offering. The total reward package will be shared with candidates during the recruitment and selection process.

Accommodations at Arm

At Arm, we want to build extraordinary teams. If you need an adjustment or an accommodation during the recruitment process, please email accommodations@arm.com. To note, by sending us the requested information, you consent to its use by Arm to arrange for appropriate accommodations. All accommodation or adjustment requests will be treated with confidentiality, and information concerning these requests will only be disclosed as necessary to provide the accommodation. Although this is not an exhaustive list, examples of support include breaks between interviews, having documents read aloud, or office accessibility. Please email us about anything we can do to accommodate you during the recruitment process.

Hybrid Working at Arm

Arm’s approach to hybrid working is designed to create a working environment that supports both high performance and personal wellbeing. We believe in bringing people together face to face to enable us to work at pace, whilst recognizing the value of flexibility. Within that framework, we empower groups/teams to determine their own hybrid working patterns, depending on the work and the team’s needs. Details of what this means for each role will be shared upon application. In some cases, the flexibility we can offer is limited by local legal, regulatory, tax, or other considerations, and where this is the case, we will collaborate with you to find the best solution. Please talk to us to find out more about what this could look like for you.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jofdav.com
Prepare application

Inside Arm

Culture, engineering, and team stories

Good distractions

Talks and stories from around this role β€” technically off-topic, practically not.

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 Β· Coffee With Developers

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet Β· LIVE

2:35 min

Preventing remote code execution in PyTorch models

BalΓ‘zs Kiss Β· World Congress 2023

1:43 min

Platform engineering as the foundation for scaling AI tools

Julia Kordick Julia Kordick Β· World Congress 2026 Europe

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White Β· World Congress 2025

1:12 min

Choosing TypeScript for complex backend applications

Maximilian Otto Maximilian Otto Β· World Congress 2024

Videos

See all

Related articles

See all