Site Reliability Engineer (SRE)

Monstro
London, UK
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Application Services Microsoft Azure Bash Shell BigQuery Cloud Computing Continuous Integration Github Intrusion Detection Systems Python (Programming Language) Log Analysis Reliability Engineering
+13 more
Runbook Data Logging Scripting Google Cloud Cloud Monitoring Apigee Kubernetes Deployment Automation Api Gateway Terraform Dynatrace Api Management Golang

Job description

Monstro is building a secure, multi-tenant platform on Google Cloud, and we’re hiring a Site Reliability Engineer to own the reliability and observability of that platform end-to-end., * Define and maintain SLOs and SLIs for our tier-1 services: API gateway, application services, identity, and edge availability

  • Build canonical dashboards and alerts in Google Cloud Monitoring, backed by structured logs and BigQuery log analytics
  • Tune alert routing so every page is actionable - kill the rest
  • Instrument services for distributed tracing and structured logging; push back on services that ship without it
  • Own error budgets and use them to prioritize reliability work over feature work when burned
  • Reduce toil: automate the top recurring page from the previous quarter
  • Maintain runbooks so every page maps to one within a cycle of first occurrence

On-call rotation and incident response

  • First responder for production alerts across monitoring, API gateway, edge defense, and CI
  • Triage severity, run the incident bridge, drive mitigation (revision rollback, traffic shift, scaling, edge block, credential rotation)
  • Own internal and external incident comms during your shift
  • Drive postmortems to closure with action items tracked as audit evidence
  • Clean written handoffs at end of shift

Our stack

  • Google Cloud Platform across multiple environments
  • Apigee X for API management
  • Cloud Run, GKE Autopilot, Cloud SQL
  • Identity Platform for customer identity
  • Cloud Armor, Cloud IDS, Security Command Center for edge and posture
  • BigQuery-backed log analytics from an org-level log sink
  • OpenTofu / Terraform for everything; GitHub Actions for CI/CD
  • Linear for work tracking

Requirements

Do you have experience in Terraform?, * Solid production experience on GCP (or comparable AWS/Azure depth with willingness to ramp on GCP fast)

  • Comfortable on-call: you’ve run incidents, written postmortems, and shipped the action items
  • Strong observability fundamentals: SLOs, log-based metrics, alert hygiene, dashboard discipline
  • Working knowledge of Kubernetes, API gateways, identity systems, and at least one IaC tool
  • Scripting / coding fluency (Python, Go, Bash) for automation and tooling
  • Good written communication - handoffs, postmortems, and runbooks are part of the job
  • Bias toward fixing the system, not the symptom

Nice to Have:

  • Apigee or another enterprise API gateway in production
  • BigQuery for log analytics or audit
  • Experience standing up observability from scratch, not just maintaining inherited dashboards
  • SOC2 or similar compliance environments

Why Join Us

You’ll be at the centre of how we bring Monstro to life for our institutional clients. Your work directly shapes the success of every implementation-getting requirements right means we deliver faster, smoother, and with fewer surprises. You’ll be joining at a foundational moment, helping to build the delivery practice from the ground up alongside a Delivery Manager who will rely on you as a critical partner from day one.

If you enjoy the puzzle of understanding complex environments, the satisfaction of a well-organised document, and the energy of working directly with clients, this is your role.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · WWC Europe 2026

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

1:33 min

Case study on adopting Kubernetes and Golang effectively

Andrew Holway · LIVE

1:06 min

Developer experience and project variety at scale

Alexandra Petri · WWC 2023

Videos

See all

Related articles

See all