> Markdown version of [/jobs/ext/13764-senior-devops-engineer](https://www.wearedevelopers.com/jobs/ext/13764-senior-devops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior DevOps Engineer - **Company:** HUMANOID - **Location:** London, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Continuous Integration, DevOps, Distributed Systems, Github, Python (Programming Language), Machine Learning, Octopus Deploy, Performance Tuning, Tensorflow, Pytorch, Multi-Cloud, Reliability of Systems, HybridCloud, Containerization, Gitlab-ci, Kubernetes, Machine Learning Operations, Hardware Infrastructure, Terraform - **Published:** May 17, 2026 - **Apply:** https://find.jobs/jobs-near-me/senior-devops-engineer-city-of-london-london/2774187237-2/ ## About the Role * 5+ years of experience in DevOps, MLOps, or infrastructure engineering (Senior/Staff level) * Strong experience with Kubernetes and containerized workloads at scale * Proven experience with Infrastructure-as-Code (Terraform, Helm, or similar) * Deep familiarity with at least one major cloud provider (AWS preferred) * Solid experience building CI/CD systems (e.g., GitHub Actions, GitLab CI, ArgoCD) * Proficiency in Python for automation and tooling * Strong understanding of distributed systems, networking, and system reliability * Ability to operate independently and drive large infrastructure initiatives Nice to have: * Hands-on experience with multi-GPU and/or distributed compute environments * Experience with GPU scheduling/orchestration (e.g., Kubernetes schedulers - Volcano, Ray, etc.) * Experience supporting ML workloads or training pipelines (PyTorch, TensorFlow, etc.) * Experience with multi-cloud or hybrid cloud environments * Background in performance optimization for training workloads * Experience in robotics, simulation, or embodied AI systems ## Description Humanoid is the first AI and robotics company in the UK, creating the world's most advanced, reliable, commercially scalable, and safe humanoid robots. Our first humanoid robot HMND 01 is a next-gen labour automation unit, providing highly efficient services across various use cases, starting with industrial applications. Our Mission At Humanoid we strive to create the world's leading, commercially scalable, safe, and advanced humanoid robots that seamlessly integrate into daily life and amplify human capacity. We are building large-scale compute infrastructure for training next-generation robotics models, including transformer-based systems like VLA. This role focuses on designing and operating multi-GPU, cross-cloud platforms that enable efficient, reliable, and scalable model training. You'll work at the intersection of DevOps, MLOps, and distributed systems, helping push the limits of real-world AI. What You'll Do: * Design, build, and operate scalable multi-GPU infrastructure across cloud environments (AWS, GCP, etc.) * Own the reliability, performance, and cost-efficiency of model training platforms * Develop and maintain infrastructure-as-code and automation for provisioning, orchestration, and lifecycle management * Build and evolve CI/CD pipelines for both infrastructure and ML training workflows * Optimize distributed training workloads (scheduling, resource utilization, observability) * Ensure high standards of reliability, scalability, security, and monitoring across systems * Collaborate with ML engineers and researchers to enable efficient experimentation and productionization * Troubleshoot complex issues across distributed systems, networking, and GPU workloads * Define and implement best practices in DevOps/MLOps for a fast-scaling environment * Document systems, architecture decisions, and operational processes ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023)