> Markdown version of [/jobs/ext/77561-site-reliability-engineer-sre](https://www.wearedevelopers.com/jobs/ext/77561-site-reliability-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer (SRE) - **Company:** Monstro - **Location:** Y, France - **Salary:** €142,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Application Services, Microsoft Azure, Bash Shell, BigQuery, Cloud Computing, Continuous Integration, Github, Intrusion Detection Systems, Python (Programming Language), Log Analysis, Reliability Engineering, Runbook, Data Logging, Scripting, Google Cloud, Cloud Monitoring, Apigee, Kubernetes, Deployment Automation, Api Gateway, Terraform, Dynatrace, Api Management, Golang - **Published:** May 28, 2026 - **Apply:** https://fr.indeed.com/viewjob?jk=7d97ce73cca73188 ## About the Role Do you have experience in Terraform?, * Solid production experience on GCP (or comparable AWS/Azure depth with willingness to ramp on GCP fast) * Comfortable on-call: you've run incidents, written postmortems, and shipped the action items * Strong observability fundamentals: SLOs, log-based metrics, alert hygiene, dashboard discipline * Working knowledge of Kubernetes, API gateways, identity systems, and at least one IaC tool * Scripting / coding fluency (Python, Go, Bash) for automation and tooling * Good written communication - handoffs, postmortems, and runbooks are part of the job * Bias toward fixing the system, not the symptoms Nice to Have: * Apigee or another enterprise API gateway in production * BigQuery for log analytics or audit * Experience standing up observability from scratch, not just maintaining inherited dashboards * SOC2 or similar compliance environments, If you're excited to contribute to a high-bar team building something meaningful, we love to hear from you! ## Description Monstro is building a secure, multi-tenant platform on Google Cloud, and we're hiring a Site Reliability Engineer to own the reliability and observability of that platform end-to-end., * Define and maintain SLOs and SLIs for our tier-1 services: API gateway, application services, identity, and edge availability * Build canonical dashboards and alerts in Google Cloud Monitoring, backed by structured logs and BigQuery log analytics * Tune alert routing so every page is actionable - kill the rest * Instrument services for distributed tracing and structured logging; push back on services that ship without it * Own error budgets and use them to prioritize reliability work over feature work when burned * Reduce toil: automate the top recurring page from the previous quarter * Maintain runbooks so every page maps to one within a cycle of first occurrence On-call rotation and incident response * First responder for production alerts across monitoring, API gateway, edge defense, and CI * Triage severity, run the incident bridge, drive mitigation (revision rollback, traffic shift, scaling, edge block, credential rotation) * Own internal and external incident comms during your shift * Drive postmortems to closure with action items tracked as audit evidence * Clean written handoffs at end of shift Our stack * Google Cloud Platform across multiple environments * Apigee X for API management * Cloud Run, GKE Autopilot, Cloud SQL * Identity Platform for customer identity * Cloud Armor, Cloud IDS, Security Command Center for edge and posture * BigQuery-backed log analytics from an org-level log sink * OpenTofu / Terraform for everything; GitHub Actions for CI/CD * Linear for work tracking ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)