> Markdown version of [/jobs/ext/2264193-staff-software-engineer-kubernetes-platform](https://www.wearedevelopers.com/jobs/ext/2264193-staff-software-engineer-kubernetes-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Software Engineer, Kubernetes Platform - **Company:** Humanloop - **Location:** London, UK - **Salary:** £57,000.0 - £73,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, C++ (Programming Language), Cloud Computing, Software Debugging, Linux, Distributed Systems, Domain Name System (DNS), Python (Programming Language), Linux Kernel, Node.Js, Service Discovery, Software Engineering, Apache Zookeeper, Graphics Processing Unit (GPU), Istio, Kubernetes, Slurm, Machine Learning Operations - **Published:** August 27, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5856707746 ## About the Role * Significant software engineering experience building and operating production distributed systems * Proficiency in at least one systems-appropriate language such as Go, Python, Rust, or C++ * Deep, hands-on Kubernetes experience beyond basic usage, including scheduler, controllers, apiserver, or operating large multi-tenant clusters * Demonstrated ability to debug complex issues across the stack, from API behavior to node- and network-level root causes * A track record of designing for reliability, correctness, and clear failure semantics in systems other engineers depend on * Strong written and verbal communication skills, with comfort building consensus with internal stakeholders * Experience with Kubernetes internals or contributions such as kube-scheduler, the scheduling framework, apiserver, etcd, client-go, controller-runtime, or similar * Experience building or operating cluster schedulers or batch systems such as Kueue, Volcano, Slurm, or in-house equivalents * Background scaling control planes or coordination systems such as etcd, ZooKeeper, Consul, or large DNS/service-mesh deployments * Familiarity with ML infrastructure such as GPUs, TPUs, or Trainium; gang scheduling; topology-aware placement; or collective networking such as NCCL * Experience with GCP and/or AWS, including GKE/EKS internals and Infrastructure as Code * Low-level systems experience such as Linux kernel tuning, cgroups, or eBPF * 12+ years of relevant industry experience, including time leading large, ambiguous infrastructure projects * Bachelors degree or an equivalent combination of education, training, and/or experience * A field relevant to the role as demonstrated through coursework, training, or professional experience ## Description * Own, operate, and extend the Kubernetes scheduler for our accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemption * Scale the Kubernetes control plane, including apiserver, etcd, and controller-manager, to support clusters far beyond typical limits, and identify the next bottleneck before it finds us * Design, build, and operate core cluster services such as service discovery that every workload in the fleet depends on * Build and maintain custom controllers, operators, and CRDs * Partner with research, training, and inference to understand workload shapes and translate requirements into platform capabilities * Collaborate with cloud providers on required features and escalations * Participate in on-call, lead incident response, and design processes such as postmortems, runbooks, and SLOs to help the team avoid repeating failures Technologies: * AI * API * AWS * Cloud * GCP * Support * Kubernetes * Linux * Network * Python * Rust * ZooKeeper * NodeJS ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [One Platform Could Not Fit Them All](https://www.wearedevelopers.com/videos/1919-one-platform-could-not-fit-them-all) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Get started with securing your cloud-native Java microservices applications](https://www.wearedevelopers.com/videos/123-get-started-with-securing-your-cloud-native-java-microservices-applications) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)