Site Reliability Engineer
Natobotics
London, UK
2 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.collegerecruiter.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source
Tech stack
Amazon Web Services
Bash Shell
Cloud Computing
Computer Programming
Computer Networks
Continuous Integration
Software Debugging
Linux
Python (Programming Language)
Reliability Engineering
Site Reliability Engineering Practices
Ansible
+16 more
Prometheus
Test Data
Data Logging
Scripting
Computer Networking Systems
Grafana
Containerization
Gitlab-ci
Kubernetes
Infrastructure Automation Frameworks
Data Management
Terraform
Splunk
Serverless Computing
Docker
Jenkins
Job description
- Automate environment lifecycle: Develop Infrastructure as Code (IaC) to automate provisioning, teardown, and configuration of test environments, integrating them with the CI/CD pipeline.
- Establish service level objectives (SLOs): Define and measure SLIs for test environments, such as availability and provisioning time.
- Monitor environment health and performance: Use observability tools like Prometheus and Grafana to track the health of test environments, identify bottlenecks, and resolve issues proactively, not reactively.
- Manage incident response: Lead the incident management process for test environment issues, conducting blameless post-mortems to understand the root causes and implement lasting fixes.
- Minimize toil: Automate manual, repetitive tasks associated with test environments to free up engineering time for more strategic work.
Strategic and cultural responsibilities
- Drive continuous improvement: Analyze environment performance data, incident reports, and post-mortems to identify opportunities for continuous improvement and innovation.
- Balance reliability and speed: Use an “error budget” for test environments. If environments are highly reliable, teams can use the budget for quicker feature development. If reliability is low, the focus shifts to improving stability.
- Instil a reliability culture: Promote a blameless culture around test environment incidents, encouraging shared ownership and collaboration between development, QA, and SRE teams.
- Capacity planning: Anticipate the future resource needs of test environments by analysing usage patterns and project forecasts. Ensure the infrastructure can scale to meet demand.
- Advance test data management: Work with Test Data Managers to ensure that test data is not only readily available but also consistent, compliant, and automatically provisioned with the environments.
Requirements
- Expertise in tooling: Proficiency with monitoring and logging tools (e.g., Prometheus, Splunk, Grafana), CI/CD platforms (e.g., Jenkins, GitLab CI), and configuration management tools (e.g., Ansible, Terraform).
- Cloud infrastructure knowledge: Deep understanding of cloud platforms like AWS, including experience with containerization technologies (Docker, Kubernetes) and serverless computing.
- Scripting and programming: Strong scripting skills in languages such as Python or Bash to automate environment management tasks.
- Systems and networking knowledge: Solid understanding of Linux systems, networking concepts, and database management.
Soft Skills
- Leadership and influence: The ability to champion SRE practices and influence technical and business stakeholders across different teams.
- Problem-solving: Strong analytical and debugging skills for investigating and resolving complex environment issues under pressure.
- Communication: Excellent communication and collaboration skills to bridge the gap between development, QA, and operations teams.
- Adaptability: A proactive and adaptable mindset to keep pace with evolving technology and development methodologies.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.collegerecruiter.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
over 2 years ago
KM
Kaleb McKelvey
The Best Software Developer Blogs to Read
over 3 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
about 2 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
19 days ago
LM
Luis Minvielle
Fully Remote Software Engineer Jobs
over 2 years ago
EM
Eli McGarvie
Find a Developer Job: 12 Best Job Sites For Developers
over 3 years ago