Staff Software Engineer, Kubernetes Platform

Humanloop
London, UK
1 day ago
Apply on www.adzuna.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
£57,000.0 - £73,000.0
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services C++ (Programming Language) Cloud Computing Software Debugging Linux Distributed Systems Domain Name System (DNS) Python (Programming Language) Linux Kernel Node.Js
+8 more
Service Discovery Software Engineering Apache Zookeeper Graphics Processing Unit (GPU) Istio Kubernetes Slurm Machine Learning Operations

Job description

  • Own, operate, and extend the Kubernetes scheduler for our accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemption
  • Scale the Kubernetes control plane, including apiserver, etcd, and controller-manager, to support clusters far beyond typical limits, and identify the next bottleneck before it finds us
  • Design, build, and operate core cluster services such as service discovery that every workload in the fleet depends on
  • Build and maintain custom controllers, operators, and CRDs
  • Partner with research, training, and inference to understand workload shapes and translate requirements into platform capabilities
  • Collaborate with cloud providers on required features and escalations
  • Participate in on-call, lead incident response, and design processes such as postmortems, runbooks, and SLOs to help the team avoid repeating failures

Technologies:

  • AI
  • API
  • AWS
  • Cloud
  • GCP
  • Support
  • Kubernetes
  • Linux
  • Network
  • Python
  • Rust
  • ZooKeeper
  • NodeJS

Requirements

  • Significant software engineering experience building and operating production distributed systems
  • Proficiency in at least one systems-appropriate language such as Go, Python, Rust, or C++
  • Deep, hands-on Kubernetes experience beyond basic usage, including scheduler, controllers, apiserver, or operating large multi-tenant clusters
  • Demonstrated ability to debug complex issues across the stack, from API behavior to node- and network-level root causes
  • A track record of designing for reliability, correctness, and clear failure semantics in systems other engineers depend on
  • Strong written and verbal communication skills, with comfort building consensus with internal stakeholders
  • Experience with Kubernetes internals or contributions such as kube-scheduler, the scheduling framework, apiserver, etcd, client-go, controller-runtime, or similar
  • Experience building or operating cluster schedulers or batch systems such as Kueue, Volcano, Slurm, or in-house equivalents
  • Background scaling control planes or coordination systems such as etcd, ZooKeeper, Consul, or large DNS/service-mesh deployments
  • Familiarity with ML infrastructure such as GPUs, TPUs, or Trainium; gang scheduling; topology-aware placement; or collective networking such as NCCL
  • Experience with GCP and/or AWS, including GKE/EKS internals and Infrastructure as Code
  • Low-level systems experience such as Linux kernel tuning, cgroups, or eBPF
  • 12+ years of relevant industry experience, including time leading large, ambiguous infrastructure projects
  • Bachelors degree or an equivalent combination of education, training, and/or experience
  • A field relevant to the role as demonstrated through coursework, training, or professional experience

Benefits & conditions

We are Anthropic, a public benefit corporation headquartered in San Francisco, building reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. Our Kubernetes Platform team runs one of the industrys largest AI compute fleets across multiple cloud providers and datacenters, and we own the control plane that keeps it operating at scale. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a collaborative office space. We also have a location-based hybrid policy requiring staff to be in one of our offices at least 25% of the time, and we sponsor visas on a case-by-case basis where possible.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:54 min

Speaker background and open source Kubernetes edge computing projects

Gaurav Gahlot Gaurav Gahlot · World Congress 2026 Europe

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · World Congress 2026 Europe

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all