Site Reliability Engineer

Natobotics
London, UK
2 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Amazon Web Services Bash Shell Cloud Computing Computer Programming Computer Networks Continuous Integration Software Debugging Linux Python (Programming Language) Reliability Engineering Site Reliability Engineering Practices Ansible
+16 more
Prometheus Test Data Data Logging Scripting Computer Networking Systems Grafana Containerization Gitlab-ci Kubernetes Infrastructure Automation Frameworks Data Management Terraform Splunk Serverless Computing Docker Jenkins

Job description

  • Automate environment lifecycle: Develop Infrastructure as Code (IaC) to automate provisioning, teardown, and configuration of test environments, integrating them with the CI/CD pipeline.
  • Establish service level objectives (SLOs): Define and measure SLIs for test environments, such as availability and provisioning time.
  • Monitor environment health and performance: Use observability tools like Prometheus and Grafana to track the health of test environments, identify bottlenecks, and resolve issues proactively, not reactively.
  • Manage incident response: Lead the incident management process for test environment issues, conducting blameless post-mortems to understand the root causes and implement lasting fixes.
  • Minimize toil: Automate manual, repetitive tasks associated with test environments to free up engineering time for more strategic work.

Strategic and cultural responsibilities

  • Drive continuous improvement: Analyze environment performance data, incident reports, and post-mortems to identify opportunities for continuous improvement and innovation.
  • Balance reliability and speed: Use an “error budget” for test environments. If environments are highly reliable, teams can use the budget for quicker feature development. If reliability is low, the focus shifts to improving stability.
  • Instil a reliability culture: Promote a blameless culture around test environment incidents, encouraging shared ownership and collaboration between development, QA, and SRE teams.
  • Capacity planning: Anticipate the future resource needs of test environments by analysing usage patterns and project forecasts. Ensure the infrastructure can scale to meet demand.
  • Advance test data management: Work with Test Data Managers to ensure that test data is not only readily available but also consistent, compliant, and automatically provisioned with the environments.

Requirements

  • Expertise in tooling: Proficiency with monitoring and logging tools (e.g., Prometheus, Splunk, Grafana), CI/CD platforms (e.g., Jenkins, GitLab CI), and configuration management tools (e.g., Ansible, Terraform).
  • Cloud infrastructure knowledge: Deep understanding of cloud platforms like AWS, including experience with containerization technologies (Docker, Kubernetes) and serverless computing.
  • Scripting and programming: Strong scripting skills in languages such as Python or Bash to automate environment management tasks.
  • Systems and networking knowledge: Solid understanding of Linux systems, networking concepts, and database management.

Soft Skills

  • Leadership and influence: The ability to champion SRE practices and influence technical and business stakeholders across different teams.
  • Problem-solving: Strong analytical and debugging skills for investigating and resolving complex environment issues under pressure.
  • Communication: Excellent communication and collaboration skills to bridge the gap between development, QA, and operations teams.
  • Adaptability: A proactive and adaptable mindset to keep pace with evolving technology and development methodologies.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all