Senior AWS Site Reliability Engineer
Spectrum IT Recruitment
Charing Cross, United Kingdom
4 days ago
Role details
Contract type
Permanent contract Employment type
Full-time (> 32 hours) Working hours
Regular working hours Languages
English Experience level
Senior Compensation
£ 70KJob location
Charing Cross, United Kingdom
Tech stack
Java
Amazon Web Services (AWS)
Amazon Web Services (AWS)
Systems Engineering
Bash
C Sharp (Programming Language)
Cloud Computing
Configuration Management
Continuous Integration
DevOps
Distributed Systems
Amazon DynamoDB
Monitoring of Systems
Python
Powershell
Reliability Engineering
Ansible
Prometheus
Datadog
CircleCI
Scripting (Bash/Python/Go/Ruby)
Google Cloud Platform
Enterprise Software Applications
System Availability
Grafana
Gitlab
Cloudformation
Containerization
Gitlab-ci
Kubernetes
Infrastructure Automation Frameworks
Cloudwatch
Puppet
Rundeck
Terraform
Splunk
Docker
Pagerduty
ELK
Jenkins
Go
Programming Languages
Microservices
Job description
- We oversee the production environment by ensuring system availability and maintaining a comprehensive perspective on overall health.
- We develop tools and software to support and streamline the management of platform infrastructure and key applications.
- We focus on enhancing the dependability, performance, and delivery speed of our software products.
- We analyse and fine-tune system performance to anticipate user demands and drive innovation.
- We provide operational support and technical oversight for several large-scale distributed applications.
- We monitor and interpret system and application metrics to fine-tune performance and troubleshoot issues effectively.
- We collaborate closely with developers to enhance service quality through thorough testing and structured release practices.
- We engage in architectural discussions, manage platform operations, and contribute to capacity forecasting.
- We design and implement automated solutions to build resilient, scalable systems.
- We maintain a strong focus on delivering new features while ensuring stability and adherence to service level goals.
Technologies:
- AWS
- Lambda
- Ansible
- Bash
- C#
- CI/CD
- Cloud
- CloudWatch
- Datadog
- DevOps
- Docker
- EC2
- ELK
- GitLab
- Grafana
- Incident Management
- Support
- Java
- Jenkins
- Kubernetes
- PagerDuty
- PowerShell
- Prometheus
- Puppet
- Python
- Splunk
- Terraform
- microservices
- Fine-tuning
More:
We deliver cutting-edge enterprise software solutions across cloud and on-premises environments, helping organisations improve customer experiences, maintain regulatory compliance, and fight fraud. Our software is trusted by businesses worldwide to enable seamless, intelligent customer interactions. This role offers hybrid working with 3 days from home, along with benefits including life insurance at 4 x annual salary, private medical insurance, a bonus scheme, an employee assistance programme, and GP online assistance, plus much more.
Requirements
- We require 3-6 years of hands-on experience in a similar role, with a strong emphasis on systems engineering, automation, and service reliability.
- We are looking for proficiency in at least one programming language such as Python, Go, Java, or C#, along with scripting skills in Bash or PowerShell.
- We need a solid grasp of cloud platforms like AWS, including how core services such as EC2, ECS, Lambda, and DynamoDB operate under reliability constraints.
- We expect practical experience with infrastructure-as-code tools such as CloudFormation or Terraform.
- We require in-depth knowledge of CI/CD principles and hands-on experience with tools such as Jenkins, GitLab CI/CD, or CircleCI.
- We need strong understanding of containerization such as Docker and Kubernetes, and microservices architecture.
- We value skilled use of observability and monitoring tools such as Prometheus, Grafana, ELK stack, or AWS CloudWatch.
- We are looking for excellent analytical and troubleshooting abilities, especially within complex distributed systems.
- We expect proven experience handling incident management and conducting blameless postmortems, including leading cross-functional teams through resolution and communication during critical outages.
- We would stand out to candidates with practical experience managing large-scale Kubernetes clusters; Kubernetes certifications are a strong bonus.
- We would stand out to candidates with hands-on familiarity with the Grafana Observability Suite, including Loki, Mimir, and Tempo.
- We would stand out to candidates with a background in administering or developing with monitoring and automation tools such as Splunk, Datadog, PagerDuty, or Rundeck.
- We would stand out to candidates with experience using configuration management platforms like Ansible, Puppet, or Chef.
- We would stand out to candidates with professional certifications in cloud DevOps, such as AWS Certified DevOps Engineer or Google Cloud Professional DevOps Engineer, or similar credentials.