> Markdown version of [/jobs/ext/2415093-lead-devops-engineer](https://www.wearedevelopers.com/jobs/ext/2415093-lead-devops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead DevOps Engineer - **Company:** Collinson - **Location:** London, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Bash Shell, Continuous Availability, Continuous Integration, DevOps, Programming Tools, Disaster Recovery, Github, Identity and Access Management, Python (Programming Language), Key Management, PCI Data Security Standards, Ansible, Security Information and Event Management, Software Vulnerability Management, Datadog, Istio, Technical Debt, Firewalls (Computer Science), Amazon Virtual Private Cloud (VPC), Kubernetes, Deployment Automation, CIS Benchmarks, Api Gateway, Terraform, Serverless Computing - **Published:** August 7, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=de98ead45735af19 ## About the Role * Deep expertise with AWS at scale: EKS (Automode), Lambda, EC2, RDS/Aurora, multi-region networking (VPC + Endpoints, Transit Gateway, Network Firewall, WAF, API Gateway), security services (IAM, KMS, Secrets Manager, SCPs, Security Hub, GuardDuty) and AI tooling (Bedrock Agentcore). * Proven ownership of a production-grade EKS platform with zero-downtime deployments - ArgoCD, Istio, blue/green and canary strategies are second nature. * Terraform expert. You've set the IaC standards for a team, not just followed them. * Automation & CI/CD - Be an authority on CI/CD pipelines to automate fast, secure and seamless deployments, with a strong emphasis on improving developer experience to automate routine tasks and improve operational efficiency. * Experience embedding AI into SLDC workflows * Hands-on DR experience: automated backup, multi-region/az failover, AWS FIS testing, and formal DR exercises with defined RTO/RPO targets. * Familiarity with security - you've worked within a compliant environment, and you're comfortable in security audits and risk conversations. You're comfortable with security tooling such as CrowdStrike and Rapid7, SIEM/SOC. * Observability ownership with Datadog (or equivalent) - you've defined SLOs, built the dashboards, and set the alerting culture. * Experience leading or mentoring a DevOps team - you can set direction, manage technical debt conversations, and grow engineers. * Comfortable with Python, Bash, Ansible, Helm, and GitHub Actions. ## Description As Lead DevOps Engineer, you'll own the design, reliability, scalability, security, and cost efficiency of our production platform - a zero-downtime, multi-region AWS environment running Kubernetes and serverless workloads and centralised AI enablement layers at scale. You'll set the technical direction for the DevOps team, driving and documenting the platform standards, practices, and tooling that champion DevEx, eases friction and keep our systems resilient and secure. This is a hands-on leadership role. You'll architect and solve complex infrastructure challenges while mentoring and growing the engineers around you. You'll own our platform SLAs, serve as the platform advocate in cross-functional decisions, and act as the core authority on resilience, cost, security, and enablement., * Platform Ownership - Own the end-to-end reliability, stability, and performance of our production platform. Enforce SLAs across availability, security posture, and resilience. Drive the measures, thresholds, and incident response standards the team operates to. * Zero-Downtime Operations - Run and maintain a multi-region, zero-downtime platform. Own deployment strategies (blue/green, canary), traffic management, and failover patterns that ensure continuous availability under change and failure. * Security & Compliance - Own the platform's security posture end to end: networking controls, IAM, KMS, secrets management, vulnerability management, and compliance against frameworks such as PCI DSS v4 and CIS Benchmarks. Work closely with security teams and act as the DevOps authority on audit and risk. * DR & Business Continuity - Participate in disaster recovery strategy, automation, execution and continuous improvement. * TCO & Cost Management - Drive the strategy on platform total cost of ownership, providing teams with clear cost-allocation visibility. Drive rightsizing, tagging strategy and architectural decisions that balance cost against reliability and developer experience. * Infrastructure as Code & CI/CD - Set and maintain the standards for IaC (Terraform, Terragrunt). Ensure deployments are secure, fast, and auditable. Raise the bar on automation so routine operational tasks are eliminated, not managed. * Observability - Own the observability strategy across the platform using Datadog. Define what good looks like: SLO/SLA dashboards, alerting thresholds, runbooks, and the feedback loops that let teams act on signals before users feel them. * Team Leadership & Mentoring - Lead and grow a team of DevOps engineers. Set technical direction, conduct design reviews, and create an environment where engineers take ownership. Actively mentor and develop less experienced engineers, helping them move from task execution to platform thinking. * AI & Automation - Lead the team's adoption of AI-assisted operations - embedding AI into CI/CD pipelines, runbook automation, incident response, and developer tooling. Define where AI adds real leverage and own its safe integration into the platform. ## Related Videos - [We adopted DevOps and are Cloud-native, Now What?](https://www.wearedevelopers.com/videos/485-we-adopted-devops-and-are-cloud-native-now-what) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Next-gen CI/CD with Gitops and Progressive Delivery](https://www.wearedevelopers.com/videos/1603-next-gen-ci-cd-with-gitops-and-progressive-delivery) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)