> Markdown version of [/jobs/ext/3589347-lead-devops-engineer](https://www.wearedevelopers.com/jobs/ext/3589347-lead-devops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead DevOps Engineer - **Company:** Collinson - **Location:** London, UK - **Experience:** Expert - **Salary:** £103,364.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Bash Shell, Continuous Availability, DevOps, Programming Tools, Disaster Recovery, Github, Identity and Access Management, Python (Programming Language), Key Management, PCI Data Security Standards, Systems Development Life Cycle, Ansible, Security Information and Event Management, Software Vulnerability Management, Datadog, Istio, Technical Debt, AgentCore, Firewalls (Computer Science), Amazon Virtual Private Cloud (VPC), Kubernetes, Deployment Automation, CIS Benchmarks, Api Gateway, Terraform, Serverless Computing - **Published:** October 5, 2026 - **Apply:** https://www.adzuna.co.uk/jobs/details/5912292456 ## About the Role * Deep expertise with AWS at scale: EKS (Automode), Lambda, EC2, RDS/Aurora, multi-region networking (VPC + Endpoints, Transit Gateway, Network Firewall, WAF, API Gateway), security services (IAM, KMS, Secrets Manager, SCPs, Security Hub, GuardDuty) and AI tooling (Bedrock Agentcore). * Proven ownership of a production-grade EKS platform with zero-downtime deployments, with ArgoCD, Istio, blue/green and canary strategies as second nature. * Terraform expertise, having set the IaC standards for a team, not just followed them. * Authority on CI/CD pipelines to automate fast, secure and seamless deployments, with a strong emphasis on improving developer experience to automate routine tasks and improve operational efficiency. * Experience embedding AI into SDLC workflows. * Hands-on DR experience: automated backup, multi-region/AZ failover, AWS FIS testing and formal DR exercises with defined RTO/RPO targets. * Familiarity with security, having worked within a compliant environment, and comfortable in security audits and risk conversations. * Comfortable with security tooling such as CrowdStrike and Rapid7, SIEM/SOC. * Observability ownership with Datadog (or equivalent), having defined SLOs, built the dashboards and set the alerting culture. * Experience leading or mentoring a DevOps team, with the ability to set direction, manage technical debt conversations and grow engineers. * Comfortable with Python, Bash, Ansible, Helm and GitHub Actions. ## Description As Lead DevOps Engineer, you will own the design, reliability, scalability, security and cost efficiency of our production platform: a zero-downtime, multi-region AWS environment running Kubernetes and serverless workloads and centralised AI enablement layers at scale. You will set the technical direction for the DevOps team, driving and documenting the platform standards, practices and tooling that champion DevEx, ease friction and keep our systems resilient and secure. This is a hands-on leadership role. You will architect and solve complex infrastructure challenges while mentoring and growing the engineers around you. You will own our platform SLAs, serve as the platform advocate in cross-functional decisions, and act as the core authority on resilience, cost, security and enablement. Key Responsibilities * Own the end-to-end reliability, stability and performance of our production platform. * Enforce SLAs across availability, security posture and resilience. * Drive the measures, thresholds and incident response standards the team operates to. * Run and maintain a multi-region, zero-downtime platform. * Own deployment strategies (blue/green, canary), traffic management and failover patterns that ensure continuous availability under change and failure. * Own the platform's security posture end to end: networking controls, IAM, KMS, secrets management, vulnerability management and compliance against frameworks such as PCI DSS v4 and CIS Benchmarks. * Work closely with security teams and act as the DevOps authority on audit and risk. * Participate in disaster recovery strategy, automation, execution and continuous improvement. * Drive the strategy on platform total cost of ownership, providing teams with clear cost-allocation visibility. * Drive rightsizing, tagging strategy and architectural decisions that balance cost against reliability and developer experience. * Set and maintain the standards for IaC (Terraform, Terragrunt). * Ensure deployments are secure, fast and auditable. * Raise the bar on automation so routine operational tasks are eliminated, not managed. * Own the observability strategy across the platform using Datadog. * Define what good looks like: SLO/SLA dashboards, alerting thresholds, runbooks and the feedback loops that let teams act on signals before users feel them. * Lead and grow a team of DevOps engineers. * Set technical direction, conduct design reviews and create an environment where engineers take ownership. * Actively mentor and develop less experienced engineers, helping them move from task execution to platform thinking. * Lead the team's adoption of AI-assisted operations, embedding AI into CI/CD pipelines, runbook automation, incident response and developer tooling. * Define where AI adds real leverage and own its safe integration into the platform.