> Markdown version of [/jobs/ext/1495808-principal-observability-cloud-platform-engineer](https://www.wearedevelopers.com/jobs/ext/1495808-principal-observability-cloud-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Observability & Cloud Platform Engineer - **Company:** 17918 - **Location:** Cambridge, UK - **Contract:** Permanent contract - **Skills:** Amazon Web Services, ARM Architecture, Cloud Computing, Distributed Computing Environment, Python (Programming Language), Open Source Technology, Prometheus, Parquet, Cloud Platform System, Istio, Grafana, Multi-Cloud, Kubernetes, Low Latency, Free and Open-Source Software, Terraform - **Published:** July 30, 2026 - **Apply:** https://www.apply4u.co.uk/jobs/x/42357002/ ## About the Role watched it run. Depth across the open-source observability stack: Prometheus, Grafana, and large-scale metrics (Thanos, Mimir, Cortex or VictoriaMetrics) logs (Loki / ELK / OpenSearch) traces (Tempo). Kubernetes at multi-cluster scale, service mesh (Istio / Envoy), Terraform, and AWS and/or GCP. A track record of evolving storage and query architectures (TSDB, Parquet, distributed processing) for cost, scale and latency. Nice to have: OpenTelemetry / OpenMetrics standards work, CNCF open-source contributions, security-in-platform experience, and using AI tooling to cut toil. TPBN1_UKTJ ## Description Principal Observability & Cloud Platform Engineer Most observability engineers run someone else's stack. This role is for the person who builds it. Our client is re-architecting observability and cloud infrastructure at a scale very few engineers ever touch: a 3,000-node Kubernetes estate, 50TB of logs a day (around 600k logs/second) and up to 80 million active time-series, running multi-region and multi-cloud across AWS and GCP. You'll own the architecture: metrics, logs, traces, telemetry pipelines, service mesh and developer experience for thousands of services and millions of devices. You'll overhaul core open-source components, storage layers, query paths for performance, cost and reliability, and push improvements back upstream to CNCF projects. This is hands-on architecture, not stack-sitting. What you'll need: Strong, hands-on Go in production, plus Python or Shell. Real scale: PB-level ingestion and hundreds of millions of active series, and you built or scaled it, not just ## Related Videos - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Parquet, Delta, Iceberg & Ducklake - An introduction for developers](https://www.wearedevelopers.com/videos/100075-parquet-delta-iceberg-ducklake-an-introduction-for-developers) - [Tracking vehicles at scale](https://www.wearedevelopers.com/videos/1999-tracking-vehicles-at-scale) - [Keycloak case study: Making users happy with service level indicators and observability](https://www.wearedevelopers.com/videos/1599-keycloak-case-study-making-users-happy-with-service-level-indicators-and-observability) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Effortlessly Scale Prometheus With The Telemetry Data Platform – And Keep your Grafana Dashboards, Too!](https://www.wearedevelopers.com/magazine/3-effortlessly-scale-prometheus-with-the-telemetry-data-platform-and-keep-your-grafana-dashboards-too) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss)