> Markdown version of [/jobs/ext/341660-sre-engineer](https://www.wearedevelopers.com/jobs/ext/341660-sre-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SRE Engineer - **Company:** GBST - **Location:** London, UK - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, Continuous Integration, Disaster Recovery, Amazon DynamoDB, Elasticsearch, Monitoring of Systems, Identity and Access Management, Apache JMeter, Python (Programming Language), Load Testing, Nginx, RabbitMQ, Reliability Engineering, Prometheus, Ruby, Strategies of Testing, Datadog, SSL Certificate Management, Data Logging, File Transfer Protocol (FTP), Autoscaling, Istio, System Availability, Grafana, Mttr, Reliability of Systems, Cloudformation, Kubernetes, Infrastructure Automation Frameworks, Linkerd (Service Mesh), Cloudwatch, Api Gateway, Terraform, New Relic (SaaS), Docker, Pagerduty - **Published:** June 14, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=516dbe98408d06c8 ## About the Role Do you have experience in Terraform?, * Ability to work on multiples tasks in parallel * Problem solver * Excellent communicator * Desire to improve things What skills you will need? * Kubernetes o Kubernetes and application troubleshooting o Application deployment GitOps / ArgoCD o K8s and application logging (Loki / fluent bit) o Service Mesh (Linkerd preferred) o Ingress Config / Troubleshooting (AWS LB Controller / Nginx) o Autoscaling configuration (Karpenter) o Certificate management (cert-manager) * AWS services o EKS o RDS, DMS, RDS Proxy o AWS Backup o API Gateway o RabbitMQ o AWS Transfer Family (SFTP / SFTP Connector) o AWS NGFW, TGW, PrivateLink o AppStream o Lambda - Python o IAM o Kinesis o DynamoDB * Terragrunt / Terraform o Troubleshooting defects * GitOps o Helm / ArgoCD * Observability Tooling o Grafana, Prometheus, Loki, Cloudwatch configuration/dashboard creation * CI/CD, * Strong knowledge of container orchestration tools like Kubernetes and Docker. * Familiarity with deploying infrastructure as Code (IaC) with Terraform and CloudFormation. * Chaos Engineering Proficiency: * Understanding of implementing resilience testing strategies * Designing and implementing chaos engineering tools like AWS Fault Injection, Gremlin, Chaos Monkey, or LitmusChaos to design and execute fault injection experiments. * Knowledge of modern chaos engineering trends, such as adaptive resilience testing or AI driven fault detection. * Monitoring and Observability: * Experience with monitoring and observability tools (e.g., Prometheus, ADOT, Grafana, Datadog, New Relic, Elastic Stack). * Strong understanding of instrumenting infrastructure with metrics, logging, and tracing * Automation and Scripting: * Proficiency in scripting and automation languages (e.g., Python, Go, Shell, Ruby, or Java). * Demonstrated ability to automate infrastructure and operational processes. * Incident Management and Root Cause Analysis: * Participating in incident response processes, including triage, mitigation, and communication. * Familiarity with incident management tools like PagerDuty or Opsgenie. * Responding to production incidents, troubleshoot issues across the full stack, and ensure minimal downtime by driving root cause analysis and applying long-term fixes. * Conducting blameless post-mortems to identify root causes and derive actionable insights, ensuring continuous improvement. * Developing playbooks for common incidents, reducing Mean Time to Resolution (MTTR) * Resilience and Scalability Design: * Understanding of system design principles, scalability, and high-availability architectures. * Practical experience with load testing and performance benchmarking tools (e.g., JMeter, Locust, k6). * Designing and testing disaster recovery (DR) strategies to ensure minimal downtime and data ## Description We're now on the lookout for a SRE Engineer. You'll be joining a global, diverse team working with cross-functional stakeholders. This is a permanent full time opportunity based in London., The type of person suitable for this role, * Managing and optimising our infrastructure to ensure high availability and system reliability. * Deliver 24/7 support via on call rotation for after hour issues * Infrastructure Automation Expertise: * Experience with the AWS cloud platform including designing, deploying, and maintaining scalable infrastructure. ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [Get started with securing your cloud-native Java microservices applications](https://www.wearedevelopers.com/videos/123-get-started-with-securing-your-cloud-native-java-microservices-applications) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023)