> Markdown version of [/jobs/ext/3288290-platform-engineer](https://www.wearedevelopers.com/jobs/ext/3288290-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Platform Engineer - **Company:** Mako - **Location:** London, UK - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Backup Devices, Cloud Engineering, Configuration Management, Computer Networks, Continuous Integration, Linux, Disaster Recovery, Distributed Data Store, Domain Name System (DNS), Network Topologies, Hypervisor, Key Management, Kernel-Based Virtual Machine, Linux System Administration, Network Control, Network Segmentation, Node.Js, Open Source Technology, Performance Tuning, Role-Based Access Control, Ansible, Prometheus, Software Engineering, Virtual Local Area Networks, AI Infrastructure, Ceph (Software), Load Balancing, CheckMK, Large Language Models, Grafana, Software Troubleshooting, Templating, Kubernetes, Bare Metal, CIS Benchmarks, Terraform - **Published:** September 24, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=76da95c0eb7173ec ## About the Role * Solid production experience running Kubernetes in a self-managed, on-premises context (not just EKS/GKE/AKS) - you understand what breaks when there's no managed control plane, and you've operated etcd and the control plane yourself * Hands-on experience designing and operating a Linux KVM/libvirt-based VM fabric as the foundation for Kubernetes - host architecture, templating, and provisioning automation, ideally from relatively early stage rather than just inheriting a mature environment * Strong Linux systems administration background (networking, storage, service management, kernel tuning, troubleshooting under pressure) * Practical, in-depth knowledge of Kubernetes networking (CNI internals, service meshes a plus) and storage (CSI drivers, distributed storage systems such as Ceph/Longhorn) * Experience with infrastructure-as-code and configuration management (Terraform, Ansible, Packer) * Experience with GitOps workflows and CI/CD pipelines * Comfortable with observability stacks (Prometheus/Grafana, ELK/Loki) and using them to diagnose infra issues without cloud-native tooling * Security-conscious: familiar with hardening standards, RBAC, network segmentation, and secrets management * Strong troubleshooting instincts across the full stack - hypervisor, OS, network, container runtime, Kubernetes control plane * Good written and verbal communication; comfortable working with distributed/hybrid teams Desirable * Experience with bare-metal Kubernetes provisioning * Background in a regulated or air-gapped/restricted-network environment * Contributions to open-source infrastructure tooling * Experience using AI tooling (e.g. AI coding assistants, LLM-based automation) to accelerate development - we're keen to use AI to speed up our development cycle * Experience running AI infrastructure. ## Description We're looking for a Platform Engineer to help design, build, and operate our on-premises Kubernetes platform running on a self-managed Linux VM fabric (KVM/libvirt-based hypervisor layer). You'll own the infrastructure that sits beneath our application workloads - from the hypervisor and VM provisioning up through Kubernetes cluster lifecycle, networking, storage, and the golden-path tooling that application teams use to ship software. This is a hands-on, deeply technical role for someone who enjoys operating infrastructure at the systems level - not a cloud-managed-service consumer role. The VM fabric itself is still being designed and built out, so you'll have real influence over its architecture, not just its day-to-day operation. You'll be responsible for keeping the fabric and the clusters running on top of it healthy, secure, and performant, without the safety net of a hyperscaler's managed control plane. What you'll be involved in: * Design the Linux VM fabric underpinning the platform from the ground up: hypervisor architecture (KVM/libvirt), host networking topology, storage backing for VM disks, and how the fabric will scale as workload demand grows * Operate and maintain the hypervisor layer across multiple physical hosts, including host patching, live migration/evacuation, and failure response with minimal workload disruption * Design and maintain a distributed shared storage system such as Ceph, or an alternative * Design and maintain VM templating and golden-image pipelines so Kubernetes nodes are provisioned consistently and can be rebuilt or rotated on demand * Automate the VM lifecycle end-to-end - provisioning, scaling, patching, decommissioning - via infrastructure-as-code * Manage compute, memory, and storage capacity planning across the fabric, including host-level oversubscription strategy and headroom for node failure or maintenance * Own virtual networking within the fabric - host networking, VLANs/overlay networks - and design how it hands off cleanly into the Kubernetes CNI layer above it * Design, build, and maintain the full lifecycle of on-prem Kubernetes clusters: bootstrapping, version upgrades, node scaling, and decommissioning, using tooling such as kubeadm, Cluster API, or Kubespray * Manage the control plane end-to-end, including etcd operations (backup/restore, performance tuning, disaster recovery), since there's no managed control plane to fall back on * Configure and tune cluster networking: CNI selection, network policy enforcement, and on-prem load balancing * Stand up and manage ingress and internal DNS for workloads across environments * Own persistent storage integration for stateful workloads via CSI drivers * Define and enforce multi-tenancy patterns across both layers - tenant isolation and resource allocation on the VM fabric (compute, storage, network) as well as namespace/resource quota strategy, RBAC, and policy enforcement (OPA/Gatekeeper or Kyverno) at the Kubernetes layer * Build and maintain GitOps-based delivery for both cluster configuration and workloads (ArgoCD or Flux), treating cluster and infrastructure state as code * Harden hosts and clusters against security baselines (CIS benchmarks for Linux and Kubernetes), manage secrets (Vault, sealed-secrets), and keep the container runtime and node OS patched * Build observability across the full stack - from hypervisor/host health up through cluster metrics and logs (we currently use Prometheus, Grafana, OpenSearch, and Checkmk; open to alternatives) - with particular focus on the capacity and failure signals a managed cloud provider would normally surface for you * Plan and execute Kubernetes version upgrades and node OS/kernel upgrades with minimal workload disruption * Design and maintain disaster recovery and backup strategy spanning both layers - VM snapshots/backups and etcd/cluster state - so the platform can be rebuilt from bare infrastructure if required * Troubleshoot incidents across the entire stack - from a misbehaving pod, down through kubelet, container runtime, and CNI, into the underlying VM and hypervisor layer when needed * Participate in an on-call rotation for platform-level incidents; drive root-cause analysis and post-incident reviews * Partner with application teams to define and support a smooth developer experience (self-service namespaces, CI/CD integration, internal developer platform tooling) * Coordinate with datacentre/facilities and network teams on physical host provisioning, rack capacity, and hardware refresh cycles * Contribute to the platform roadmap: capacity growth, tooling upgrades, and reducing operational toil through automation ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [How I saved 200K/yr in direct costs writing 0 code lines in K8s](https://www.wearedevelopers.com/videos/1055-how-i-saved-200k-yr-in-direct-costs-writing-0-code-lines-in-k8s) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path)