Principal Observability & Cloud Platform Engineer

17918
Cambridge, UK
12 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Amazon Web Services ARM Architecture Cloud Computing Distributed Computing Environment Python (Programming Language) Open Source Technology Prometheus Parquet Cloud Platform System Istio Grafana Multi-Cloud
+4 more
Kubernetes Low Latency Free and Open-Source Software Terraform

Job description

Principal Observability & Cloud Platform Engineer Most observability engineers run someone else’s stack. This role is for the person who builds it. Our client is re-architecting observability and cloud infrastructure at a scale very few engineers ever touch: a 3,000-node Kubernetes estate, 50TB of logs a day (around 600k logs/second) and up to 80 million active time-series, running multi-region and multi-cloud across AWS and GCP. You’ll own the architecture: metrics, logs, traces, telemetry pipelines, service mesh and developer experience for thousands of services and millions of devices. You’ll overhaul core open-source components, storage layers, query paths for performance, cost and reliability, and push improvements back upstream to CNCF projects. This is hands-on architecture, not stack-sitting. What you’ll need: Strong, hands-on Go in production, plus Python or Shell. Real scale: PB-level ingestion and hundreds of millions of active series, and you built or scaled it, not just

Requirements

watched it run. Depth across the open-source observability stack: Prometheus, Grafana, and large-scale metrics (Thanos, Mimir, Cortex or VictoriaMetrics) logs (Loki / ELK / OpenSearch) traces (Tempo). Kubernetes at multi-cluster scale, service mesh (Istio / Envoy), Terraform, and AWS and/or GCP. A track record of evolving storage and query architectures (TSDB, Parquet, distributed processing) for cost, scale and latency. Nice to have: OpenTelemetry / OpenMetrics standards work, CNCF open-source contributions, security-in-platform experience, and using AI tooling to cut toil. TPBN1_UKTJ

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.apply4u.co.uk

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · WWC Europe 2026

2:50 min

How Parquet metadata enables efficient data reading

Matthias Niehoff Matthias Niehoff · WWC Europe 2026

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

4:11 min

Building a unified stack leveraging core open source tools

Mathias Palmersheim Mathias Palmersheim · Europe 2026 Virtual

1:54 min

Speaker background and open source Kubernetes edge computing projects

Gaurav Gahlot Gaurav Gahlot · WWC Europe 2026

Videos

See all

Related articles

See all