Remote Senior Site Reliability Engineer

Xtremepush
Exeter, UK
3 months ago
Apply on find.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

PHP (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Big Data Software as a Service Cloud Computing Databases Continuous Integration Linux DevOps Distributed Systems
+13 more
Python (Programming Language) MySQL Performance Tuning Reliability Engineering Prometheus Caching Vue.js Event Driven Architecture Data Analytics Apache Kafka Vertica Terraform Golang

Job description

  • Act as a senior member of the SRE team, supporting activities including the backlog and workload of the team, scoping requirements, peer review of code, providing feedback to the rest of the team.
  • Represent the team in management and stakeholder meetings. Ensure best practices are kept, and suggest improvements to our development processes where you see gaps.
  • Investigate, test, and resolve technical problems, working closely with other engineers to deliver core product functionality.
  • Defining SLOs, SLIs, and SLAs for key metrics that indicate the health, security, stability and uptime of production, staging and development environments
  • Monitoring the above environments and reacting to alerts and issues that may arise in day-to-day operation of their product line.
  • Participate in an on-call rota for priority-1 level alarms with the rest of the Platform teams
  • Ongoing upgrades and improvements to operational processes to optimise performance, stability and cost.
  • Working with the platform engineering team to contribute to the planning of how we carry application/infrastructure releases and configuration changes.
  • Interact with internal teams and external 3rd party vendors to troubleshoot and resolve complex problems, * 5+ years experience in an engineering role responsible for supporting a scaled SaaS platform running on Linux in a cloud environment

Requirements

We are seeking a Senior SRE with experience of working with scaled SaaS production infrastructure. The successful candidate will work as part of a team focused on site reliability, security, and scalability, as we manage our rapid growth.

The ideal candidate will be a proactive and driven individual, who excels at understanding and working on complex technical solutions requiring performance and optimisation at scale. Our core technologies include PHP, MySQL, Vue.js and AWS. Participating in an on-call roster is required as part of this role., * Experience working with high-performance systems, and solving complex engineering problems at scale (our platform processes ~100 Billion messages per year)

  • Understanding of distributed systems design - including asynchronous tasks, event driven architecture, scheduling, caching and queue processing
  • Ability to apply distributed systems design knowledge to resolve scaling constraints. The capability to carry out performance tuning from the API to Application to Database layer of the platform.
  • Strong communication skills and ability to explain complex technical solutions simply to others
  • Strong understanding of PHP, GoLang, MySQL, Opentelemetry, Prometheus
  • Experience with Cloud and DevOps technologies (AWS, Terraform, CI/CD etc.)
  • Experience with specific technologies in our stack: Clickhouse, Kafka, Pulsar, Python
  • Experience with networking and security concepts
  • Interest or experience with marketing technologies
  • Interest or experience with big data, data analytics, AI and machine learning

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on find.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:18 min

Scaling MySQL databases for massive user growth

Johannes Nicolai Johannes Nicolai +1 · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

8:22 min

Simulating a Linux terminal and running Spring Boot

Jakov Semenski · LIVE

1:48 min

Analyzing network packets with database protocol tools

Daniël van Eeden Daniël van Eeden · World Congress 2026 Europe

Videos

See all

Related articles

See all