> Markdown version of [/jobs/ext/2305222-site-reliability-engineer-spacetime-uk](https://www.wearedevelopers.com/jobs/ext/2305222-site-reliability-engineer-spacetime-uk). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer- Spacetime UK - **Company:** Aalyria - **Location:** Greater London, UK - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Amazon Web Services, C++ (Programming Language), Software Debugging, Distributed Systems, Java Virtual Machine (JVM), Java Web Services, Python (Programming Language), Performance Tuning, Reliability Engineering, Prometheus, System Programming, Systems Integration, Data Logging, Computer Networking Systems, Google Cloud, Istio, System Availability, Grafana, Multi-Cloud, Infrastructure as Code (IaC), SC Clearance, Gitlab-ci, Kubernetes, Linkerd (Service Mesh), Terraform, Dynatrace - **Published:** August 30, 2026 - **Apply:** https://www.collegerecruiter.com/job/2815109245-site-reliability-engineer-spacetime-uk ## About the Role This is a greenfield/brownfield opportunity. You will be a trusted expert, helping to define and implement the strategy and building the tools that empower our engineers. You will support the roadmap to mature our observability stack, moving from cloud-native tools to a robust, scalable, and insightful platform built on best-in-class technologies (Prometheus, OpenTelemetry, etc.). If you are an SRE who thrives on platform-building challenges and wants to be relied upon to build a production-grade observability stack from the ground up, this role is for you., * 4+ years of experience in an SRE or platform engineering role, with a focus on observability for large-scale, distributed compute or network systems. * Deep, hands-on expertise building, scaling, and managing observability platforms (e.g., Prometheus, Grafana, Loki/ELK, OpenTelemetry, Tempo/Jaeger, Honeycomb, etc.). You have proven experience using these tools to support performance analysis and debugging of complex distributed systems. * Strong production-level experience with Google Cloud Platform (GCP) and Kubernetes. * Experience using Infrastructure as Code (IaC) and GitOps principles (e.g., ArgoCD). * Proficiency in a systems programming language, with a strong preference for Go and Python for debugging and writing tooling. * Demonstrable experience defining, implementing, and managing SLOs, SLIs, and error budgets for production services for high availability distributed systems., * Experience operating a multi-cloud environment, specifically GCP and AWS. * Hands-on experience with GitLab CI for CI/CD pipelines. * Working knowledge of service mesh technologies such as Istio or Linkerd. * Familiarity with instrumenting applications written in Go and C++. * An active Secret clearance, or higher, is preferred for this position. * Experience with JVM observability (tuning, monitoring) for Java-based applications. ## Description * Help design and build Aalyria's centralized observability platform, integrating and scaling tools for metrics (e.g. Prometheus), logging (e.g. Loki), and distributed tracing (e.g. Tempo/OpenTelemetry). * Define, implement, and manage a robust framework of Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for our core products, ensuring we are launch-ready. * Partner with SWEs to implement observability best practices, develop standard templates and documentation, and configure tooling (e.g., OpenTelemetry libraries). * Automate the deployment, scaling, and management of the entire observability stack using Infrastructure as Code (e.g. Terraform) and GitOps principles (e.g. ArgoCD). * Partner closely with the core infrastructure team to ensure deep visibility into our Kubernetes clusters and underlying GCP and AWS environments. * Develop and lead the company's monitoring, alerting, and incident response strategy, driving a culture of proactive reliability and blameless post-mortems. ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [The Power of Purpose: Unlocking Potential and Innovation](https://www.wearedevelopers.com/videos/1110-the-power-of-purpose-unlocking-potential-and-innovation) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)