Senior DevOps / SRE
Experis
Charing Cross, United Kingdom
yesterday
Role details
Contract type
Permanent contract Employment type
Full-time (> 32 hours) Working hours
Regular working hours Languages
English Experience level
SeniorJob location
Charing Cross, United Kingdom
Tech stack
Artificial Intelligence
Azure
Cloud Computing
Cloud Engineering
Continuous Integration
DevOps
Python
Machine Learning
Reliability Engineering
Azure
Google Cloud Platform
Spring Cloud
Large Language Models
Software Troubleshooting
Reliability of Systems
Kubernetes
Free and Open-Source Software
Terraform
Docker
Service Stack
Job description
This is an opportunity for an engineer who thrives in production environments, enjoys solving complex operational challenges, and is passionate about building resilient, highly available platforms. You'll play a key role in maintaining and improving cloud infrastructure, automation, observability, and system reliability., * Support and operate cloud-native applications within Google Cloud Platform (GCP)
- Build and maintain reliable, scalable production systems with a strong SRE focus
- Monitor, troubleshoot, and resolve production issues while driving continuous reliability improvements
- Work with Python to automate operational tasks, tooling, integrations, and platform enhancements
- Manage Kubernetes environments and containerised workloads
- Develop and maintain CI/CD pipelines to support efficient and reliable deployments
- Manage infrastructure through Infrastructure as Code using Terraform
- Implement best practices around observability, monitoring, incident response, and platform resilience
- Collaborate with engineering teams to improve operational performance and system reliability
Requirements
- Strong Google Cloud Platform (GCP) experience
- Proven background in Site Reliability Engineering (SRE), DevOps, or cloud operations
- Strong Python skills, with experience using Python for automation, tooling, and operational engineering
- Hands-on experience with Kubernetes, Terraform, Docker, and CI/CD pipelines
- Experience supporting and operating production systems at scale
- Strong troubleshooting, incident management, and root-cause analysis skills
- Understanding of reliability, scalability, observability, and performance engineering
- A collaborative approach with a strong sense of ownership and continuous improvement
Desirable Skills
- Experience with Microsoft Azure
- Exposure to AI/ML platforms or machine learning workloads
- Experience supporting LLM, NLP, or agent-based systems
- Experience with multi-container architectures
- Open-source contributions or technical mentoring experience
- Experience within highly regulated or scientific environments
What's on Offer?
- Opportunity to work on innovative AI and machine learning technologies
- High levels of ownership and technical autonomy
- Modern cloud-native technology stack
- Collaborative, engineering-led culture
- The chance to make a meaningful impact through reliable, scalable technology
If you're an experienced SRE with strong Python and GCP expertise and enjoy improving the reliability and performance of complex cloud environments, we'd love to hear from you.