Site Reliability Engineer

iXceed Solutions
Basildon, UK
1 day ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Amazon Web Services Microsoft Azure Backup Devices Bash Shell Cloud Computing Cloud Engineering Computer Programming Computer Networks Continuous Integration DevOps Disaster Recovery Python (Programming Language)
+23 more
Octopus Deploy Performance Tuning Reliability Engineering Prometheus Software Deployment Systems Architecture Data Logging Istio Grafana Reliability of Systems HybridCloud Cloudformation Containerization Gitlab-ci Kubernetes Infrastructure Automation Frameworks Information Technology Performance Monitor Linkerd (Service Mesh) Api Gateway Terraform Jenkins Golang

Job description

The SRE will be responsible for designing, deploying, automating, and supporting highly available, scalable, and secure containerized applications in cloud-native environments. You will work closely with development, operations, and security teams to ensure the reliability, performance, and efficiency of our production systems., * Design, deploy, and manage Kubernetes clusters on-premises and/or cloud-managed such as EKS, AKS, GKE to support scalable microservices architectures.

  • Automate infrastructure provisioning and application deployment using Infrastructure as Code (IaC) tools such as Terraform, Helm, or CloudFormation.
  • Monitor, troubleshoot, and optimize system performance using observability tools.
  • Implement and manage CI/CD pipelines to ensure rapid, repeatable, and reliable software delivery.
  • Ensure system reliability, availability, and security through proactive monitoring, incident response, and root cause analysis.
  • Develop and maintain runbooks, dashboards, and documentation for operational procedures and system architectures.
  • Participate in on-call rotations and respond to production incidents, ensuring minimal downtime and fast recovery.
  • Collaborate with development and operations teams to drive DevOps and SRE best practices including capacity planning, scaling, and cost optimization.
  • Continuously improve automation tooling and processes to reduce manual work and increase system reliability.

Requirements

We are seeking a highly skilled Site Reliability Engineer (SRE) with deep expertise in Kubernetes and cloud technologies AWS, Azure, or GCP., * 3 years experience as an SRE, DevOps Engineer, or similar role supporting large-scale systems.

  • Expertise in Kubernetes deployment, scaling, upgrades, troubleshooting, and networking.
  • Hands-on experience with at least one major cloud provider (AWS, Azure, or GCP).
  • Proficiency in scripting/programming (Python, Bash, Go, etc.).
  • Experience with IaC tools (Terraform, Helm, CloudFormation, ARM, etc.).
  • Strong knowledge of Linux systems administration and networking concepts.
  • Familiarity with monitoring, logging, and ingestion tools (Prometheus, Grafana, ELK, EFK).
  • Experience with CI/CD tools (Jenkins, GitLab CI, ArgoCD, etc.).
  • Understanding of security best practices in cloud and containerized environments.
  • Excellent troubleshooting and problem-solving skills.
  • Strong communication and collaboration skills., * Certified Kubernetes Administrator (CKA) or similar certification.
  • Experience with service mesh (Istio, Linkerd), ingress controllers, and API gateways.
  • Experience in a multicloud or hybrid cloud environment.
  • Familiarity with GitOps practices and tools (ArgoCD, Flux).
  • Experience with disaster recovery, backup, and business continuity planning., Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.

This role is ideal for engineers who are passionate about automation, reliability, and modern cloud-native architectures and who thrive in fast-paced, collaborative environments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · World Congress 2026 Europe

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

7:15 min

Installing Istio programmatically with bash scripts

Thomas Südbröcker · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all