Lead DevOps Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+1 more
Job description
Reporting to the Head of DevOps, you’ll provide technical leadership across cloud infrastructure, platform engineering and automation while staying hands-on in delivery. The Head of DevOps owns the overall function and people leadership across our UK and Pune capabilities. You’ll be the senior technical leader, working alongside the engineers.
You won’t lead from the sidelines. You’ll be in the infrastructure, pipelines and code alongside the team, shaping our DevOps practice and turning good engineering principles into solutions that work in production.
You’ll design and evolve secure, scalable and cost-effective infrastructure, keeping our platforms reliable, observable and efficient. Working with engineering, security and product teams, you’ll take difficult technical challenges from design through to a stable production outcome.
If you want real technical depth with genuine leadership, where you can set direction and still know what’s happening under the hood, this is built for you.
What You’ll Do
- Own our AWS and Azure infrastructure. You’ll set practical standards for security, reliability, scalability, performance, compliance, cost and resource use, working with architects, developers, security and product teams to ensure they work in the real world.
- Write and maintain the Terraform that runs our infrastructure. Build reusable modules, manage state and environments, and keep documentation up to date. Use the consoles to diagnose problems, but make lasting changes through code, review and controlled deployment.
- Get hands-on with our CircleCI pipelines, creating standards practical enough to be adopted, not simply written down. Embed secret detection, dependency and container scanning, software bills of materials, code-quality checks and supply-chain controls, then help teams fix the risks that matter.
- Run our production Kubernetes platforms day-to-day, working directly with Argo CD, Helm and container tooling to deploy, upgrade, optimise and troubleshoot.
- Keep a close eye on availability, performance and resource consumption. Improve metrics, logs, dashboards and alerts so the team can spot trouble coming. When something goes wrong, lead the technical response, root-cause analysis and preventative actions.
- Ensure we can recover as effectively as we can run. Help maintain and test our backup, resilience and disaster-recovery arrangements, using Python or Bash to remove repetitive work and improve consistency.
- Mentor engineers through pairing, code reviews and hands-on problem-solving. Spread knowledge, share responsibility and avoid creating reliance on a single person.
- Stay curious. Run technical spikes and bring evidence when you think we should change direction. Be the practical bridge between engineering, security, product and the wider business, grounded in what you’re building and running.
Requirements
- You’ve run production cloud infrastructure in practice, with deep expertise in AWS or Azure and confidence in the other. You’re comfortable in the console and in the code, and you understand networking, identity, security controls and governance.
- You’ve built reusable Terraform module catalogues, managed remote state, and supported multiple environments, accounts, or subscriptions through reviewed and controlled changes.
- Your Kubernetes experience is production-focused: deployments, upgrades, troubleshooting, and capacity and performance. You’ve used Helm and GitOps day-to-day, with direct Argo CD experience strongly preferred.
- You’ve designed CI/CD pipelines from scratch and implemented DevSecOps controls yourself, including secrets detection, dependency and container scanning, SBOM generation, and software supply-chain security.
- You use Python or Bash regularly and can demonstrate tooling you’ve developed and maintained. You understand observability, incident response, root-cause analysis and production reliability.
- You’ve helped engineers grow through collaboration and shared problem-solving. You lead by example and clearly explain technical risks, options and recommendations to both technical and non-technical audiences.
Desirable Experience
- Terragrunt.
- CircleCI.
- Cloud cost optimisation and FinOps practices.
- Designing and testing backup, resilience and disaster-recovery solutions.
- Working within a geographically distributed engineering team.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
DevOps Engineer Salary [2023]
What Are The Top Skills Required For Azure Developers?
Fully Remote Software Engineer Jobs
Dev Digest 120 - Apple and peers