Site Reliability Engineer (SRE)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+7 more
Requirements
Job DescriptionJob Title: Site Reliability Engineer (SRE) - ObservabilityLocation: RemoteAbout the RoleWe are looking for a Lead SRE to design, scale, and operate massive-scale observability systems that keep our global services online and performant. You will join an autonomous team of software engineers focused on solving complex data infrastructure challenges.Key ResponsibilitiesScale Prometheus metrics infrastructure to handle 100+ million active series .Operate large Elasticsearch clusters holding 2000+TB of data .Grow high-throughput Kafka data pipelines processing hundreds of thousands of events per second.Build custom alerting workflows and self-service APIs for internal engineering teams.Provision cloud and private infrastructure using Terraform .Requirements5+ years operating mid-to-large distributed systems on Linux VMs or bare-metal machines.2+ years developing in Go, Python, Ruby, Scala, or Bash.Hands-on experience with Prometheus/Thanos/Cortex, Kafka, the ELK stack, Ansible, or Consul .Comfortable diving into unfamiliar codebases and participating in an on-call rotation.Keywords: Observability, Monitoring, SRE, Site Reliability Engineering, DevOps, ElasticSearch, ELK, Prometheus, Kafka, Terraform, Linux, Bare MetalRandstad Technologies is acting as an Employment Business in relation to this vacancy.TPBN1_UKTJ
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.apply4u.co.ukGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Fully Remote Software Engineer Jobs
Find a Developer Job: 12 Best Job Sites For Developers
Data Engineer Salary UK
Where To Find Software Engineering Jobs