Senior DevOps Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+36 more
Job description
This is a hands-on senior engineering position for someone with deep expertise across AWS, Kubernetes, infrastructure as code, CI/CD and cloud security.
You’ll be responsible for designing, building and maintaining highly available infrastructure, automating software delivery, strengthening cloud security and helping establish the operational processes required to support a growing technology organisation.
You’ll have genuine ownership of the platform and the opportunity to influence architecture, engineering standards and DevSecOps practices across the wider team. What You’ll Be Doing Cloud & Infrastructure
- Architect, implement and maintain multi-account, multi-region AWS environments.
- Build and manage infrastructure using Terraform, Terraform Cloud, Helm and CloudFormation.
- Design scalable, secure and highly available cloud-native infrastructure.
- Manage container registries and supporting platform services.
- Maintain and improve on-premises infrastructure where required.
Kubernetes & Platform Engineering
- Administer and optimise Kubernetes environments, including AWS EKS and OpenShift.
- Manage Helm-based application deployments and platform configuration.
- Support service-mesh technologies and distributed application architectures.
- Build robust container platforms using Docker and related technologies.
- Develop automation to reduce operational overhead and improve platform reliability.
CI/CD & Automation
- Design and maintain modern CI/CD pipelines using GitHub Actions and Argo CD.
- Enable automated, repeatable and low-risk deployments.
- Implement GitOps practices where appropriate.
- Automate repetitive operational processes using Bash, Python or Go.
- Support zero-touch deployment and infrastructure provisioning.
Security & DevSecOps
- Lead cloud security and infrastructure posture reviews.
- Implement and maintain AWS security best practices across IAM, networking and workloads.
- Work with services such as GuardDuty, CloudWatch and IAM Identity Center.
- Integrate SAST and DAST practices into development and deployment pipelines.
- Maintain secure software supply chains, including AMI and container image scanning.
- Implement and monitor infrastructure and application security controls.
- Support the organisation’s ongoing compliance and certification requirements.
Reliability & SRE
- Establish and maintain observability across distributed systems.
- Develop proactive monitoring, alerting and performance-tuning strategies.
- Help maintain service-level objectives and platform availability.
- Investigate and resolve infrastructure and application incidents.
- Coordinate emergency changes, hot fixes and rollbacks.
- Lead post-incident reviews and implement improvements to prevent recurrence.
Backup, Disaster Recovery & Resilience
- Design and maintain AWS Backup strategies.
- Establish and test disaster recovery processes.
- Run regular restore and recovery exercises.
- Maintain appropriate RPO/RTO reporting and dashboards.
- Ensure critical infrastructure can be recovered reliably when required.
Cost & Capacity Management
- Monitor infrastructure utilisation, availability and capacity.
- Lead regular capacity and availability reviews.
- Identify opportunities for AWS cost optimisation.
- Make effective use of AWS Compute Optimizer, Savings Plans and Graviton.
- Balance performance, availability, security and cost when making infrastructure decisions.
Governance & Compliance
- Support regular technical and security audits.
- Maintain infrastructure provenance and configuration records.
- Implement infrastructure drift detection across Terraform and Helm.
- Help ensure engineering practices remain aligned with relevant industry standards.
- Work with development teams to embed DevSecOps practices throughout the software lifecycle., * Supporting and mentoring other engineers.
- Working across multiple technical areas when required.
- Communicating clearly with technical and non-technical stakeholders.
- Working directly with clients when necessary.
- Balancing speed of delivery with security, reliability and long-term maintainability.
- Taking responsibility for systems in production rather than simply handing them over.
We’re looking for someone who builds rather than simply operates, enjoys solving difficult infrastructure problems and takes pride in leaving platforms better than they found them. Why This Role?
This is an opportunity to work on infrastructure that sits beneath frontier AI and robotics systems operating in demanding real-world environments.
You’ll have genuine ownership of the platforms you build and the opportunity to influence how the organisation approaches cloud infrastructure, security, reliability and DevOps as it scales.
You’ll be able to:
- Shape cloud and platform architecture.
- Build and harden infrastructure supporting advanced AI systems.
- Establish DevSecOps and SRE practices.
- Work across AWS, Kubernetes, automation and security.
- Develop towards technical and design leadership.
- Have a direct impact within a growing technology organisation.
Requirements
Security Clearance: UK security clearance eligibility required (British Passport + 5 years of continuous residency in the UK), We’re looking for a seasoned DevOps or SRE professional who combines strong hands-on technical skills with a pragmatic, collaborative approach.
You’ll ideally have:
- Proven professional experience in DevOps, SRE or Cloud Engineering, with recent hands-on AWS experience.
- Expert knowledge of core AWS services including EC2, VPC, IAM, S3, ALB/ELB, CloudFront, ECR/ECS and Control Tower.
- Strong experience with AWS security tooling, including GuardDuty, IAM Identity Center and CloudWatch.
- Significant experience administering Kubernetes, particularly EKS and/or OpenShift.
- Strong knowledge of Helm and Kubernetes application deployment.
- Experience building and administering on-premises Kubernetes environments.
- Experience with Nexus or comparable enterprise artifact repositories.
- Strong Infrastructure as Code experience with Terraform, Terraform Cloud and/or CloudFormation.
- Strong containerisation experience using Docker and Docker Compose.
- Excellent Linux systems administration skills.
- Strong CI/CD experience with GitHub Actions and Argo CD.
- Experience implementing SAST and DAST within software delivery pipelines.
- Strong scripting and automation skills using Bash, Python and/or Go.
- Experience delivering secure, highly available and cost-efficient cloud platforms.
- Experience implementing SSO across enterprise and industry-standard software platforms.
- A strong understanding of infrastructure security, reliability and operational best practice.
Desirable Experience
The following would be advantageous:
- MLOps or LLMOps experience.
- Experience with platforms such as SageMaker, Kubeflow or ZenML.
- Extensive on-premises Kubernetes deployment experience.
- Prometheus or comparable observability platforms.
- AWS Karpenter.
- AWS Compute Optimizer.
- Experience operating highly distributed systems.
- Familiarity with ISO 27001, NIST SSDF, OWASP SAMM or similar security frameworks.
- Understanding of GDPR fundamentals.
- Experience working within defence, national security, critical infrastructure or other regulated environments.
Security Clearance
Due to the nature of the work, applicants must be eligible to obtain UK security clearance.
Applicants will generally need to:
- Hold a UK passport.
- Have 5 years of continuous UK residency.
- Be eligible for BPSS and SC clearance., This is a small, highly collaborative engineering team, so attitude and communication are just as important as technical ability., * Taking ownership and working independently., If you’re a senior DevOps, SRE or cloud engineer who enjoys building secure, reliable infrastructure and wants genuine ownership of the platforms supporting advanced AI systems, we’d like to hear from you.
Please apply with an up-to-date CV highlighting your experience across AWS, Kubernetes, Terraform, CI/CD, DevSecOps and cloud security, together with details of any existing UK security clearance.
Benefits & conditions
- Competitive salary.
- Genuine scope and a clear path towards technical/design leadership.
- Flexible and hybrid working.
- Company pension with employer contribution.
- Private healthcare.
- 29 days’ annual leave plus public holidays.
- Opportunity to work on advanced AI, robotics and cloud infrastructure.
- Collaborative environment with significant technical ownership.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
Find a Developer Job: 12 Best Job Sites For Developers
Best Companies to work for in London: Top 25 Companies in 2023
Where To Find Software Engineering Jobs