Site Reliability Engineer- Spacetime UK

Aalyria
Greater London, UK
7 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours

Tech stack

Amazon Web Services C++ (Programming Language) Software Debugging Distributed Systems Java Virtual Machine (JVM) Java Web Services Python (Programming Language) Performance Tuning Reliability Engineering Prometheus System Programming Systems Integration
+14 more
Data Logging Computer Networking Systems Google Cloud Istio System Availability Grafana Multi-Cloud Infrastructure as Code (IaC) SC Clearance Gitlab-ci Kubernetes Linkerd (Service Mesh) Terraform Dynatrace

Job description

  • Help design and build Aalyria’s centralized observability platform, integrating and scaling tools for metrics (e.g. Prometheus), logging (e.g. Loki), and distributed tracing (e.g. Tempo/OpenTelemetry).
  • Define, implement, and manage a robust framework of Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for our core products, ensuring we are launch-ready.
  • Partner with SWEs to implement observability best practices, develop standard templates and documentation, and configure tooling (e.g., OpenTelemetry libraries).
  • Automate the deployment, scaling, and management of the entire observability stack using Infrastructure as Code (e.g. Terraform) and GitOps principles (e.g. ArgoCD).
  • Partner closely with the core infrastructure team to ensure deep visibility into our Kubernetes clusters and underlying GCP and AWS environments.
  • Develop and lead the company’s monitoring, alerting, and incident response strategy, driving a culture of proactive reliability and blameless post-mortems.

Requirements

This is a greenfield/brownfield opportunity. You will be a trusted expert, helping to define and implement the strategy and building the tools that empower our engineers. You will support the roadmap to mature our observability stack, moving from cloud-native tools to a robust, scalable, and insightful platform built on best-in-class technologies (Prometheus, OpenTelemetry, etc.). If you are an SRE who thrives on platform-building challenges and wants to be relied upon to build a production-grade observability stack from the ground up, this role is for you., * 4+ years of experience in an SRE or platform engineering role, with a focus on observability for large-scale, distributed compute or network systems.

  • Deep, hands-on expertise building, scaling, and managing observability platforms (e.g., Prometheus, Grafana, Loki/ELK, OpenTelemetry, Tempo/Jaeger, Honeycomb, etc.). You have proven experience using these tools to support performance analysis and debugging of complex distributed systems.
  • Strong production-level experience with Google Cloud Platform (GCP) and Kubernetes.
  • Experience using Infrastructure as Code (IaC) and GitOps principles (e.g., ArgoCD).
  • Proficiency in a systems programming language, with a strong preference for Go and Python for debugging and writing tooling.
  • Demonstrable experience defining, implementing, and managing SLOs, SLIs, and error budgets for production services for high availability distributed systems., * Experience operating a multi-cloud environment, specifically GCP and AWS.
  • Hands-on experience with GitLab CI for CI/CD pipelines.
  • Working knowledge of service mesh technologies such as Istio or Linkerd.
  • Familiarity with instrumenting applications written in Go and C++.
  • An active Secret clearance, or higher, is preferred for this position.
  • Experience with JVM observability (tuning, monitoring) for Java-based applications.

Benefits & conditions

  • Competitive salary benchmarked to UK aerospace and defence technology market rates.
  • Equity participation - share in Aalyria’s growth at an early stage.
  • Comprehensive benefits including pension, private health insurance, and generous annual leave.
  • Flexible and hybrid working arrangements.
  • The opportunity to work on genuinely novel technology with real-world operational impact across national security, commercial satellite, and deep-space programmes.
  • A collaborative, low-hierarchy team environment with direct exposure to technical leadership and customers.

About the company

Aalyria is a leading technology company that supplies laser communications technology and temporospatial software-defined networking platforms to the aerospace industry. With technology acquired from Google, Aalyria is at the forefront of innovation in satellite and airborne mesh networks, as well as cislunar and deep-space communications. We are revolutionizing the orchestration and management of planetary mesh networks using any radio or optical spectrum, any orbit, and any hardware across land, sea, air, and space.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · World Congress 2025

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · World Congress 2026 Europe

12:08 min

Comparing Keptn orchestration capabilities against alternative software operators

Thomas Schütz · LIVE

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all