> Markdown version of [/jobs/ext/2245217-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2245217-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Búsqueda Avanzada - **Location:** Madrid, Spain (Remote available) - **Contract:** Permanent contract - **Skills:** Configuration Management, Continuous Integration, Data Stores, Software Debugging, DevOps, Distributed Systems, Python (Programming Language), Reliability Engineering, Cloud Services, Ansible, Prometheus, Grafana, Build Management, Gitlab-ci, Kubernetes, Low Latency, Sentry, Bare Metal, Vertica, Asynchronous Programming, Jenkins - **Published:** August 26, 2026 - **Apply:** https://www.adzuna.es/contact-us.html ## About the Role + Substantial production engineering or SRE experience, including experience defining and implementing an SLO framework rather than simply operating within an existing one. + Strong Python skills and the ability to read and modify Go or Rust code when implementing instrumentation and reliability improvements. + Strong hands-on experience with time-series and event telemetry at scale, including Prometheus/OpenMetrics, Grafana, Alertmanager-class routing systems, and columnar or high-cardinality data stores such as ClickHouse or equivalent technologies. + Experience debugging distributed systems running on bare metal and long-lived hosts; this role requires more than Kubernetes-centric operational experience. + Practical experience with production-scale configuration management and CI/CD tooling such as Ansible, GitLab CI, Jenkins, or comparable technologies. + Strong understanding of telemetry for systems that cannot be directly scraped or fully controlled, including push-based collection, sampling, clock skew, partial reporting, and privacy considerations on customer-managed infrastructure. + Excellent written and asynchronous communication skills, with the ability to align multiple engineering teams around measurable definitions of system health. + Strong engineering judgment and a pragmatic approach to observability, reliability, alerting, and operational ownership. + Experience with security products such as WAF, EDR, antivirus, or vulnerability-management platforms is a strong advantage, particularly an understanding that security-control reliability must measure effective enforcement rather than simple uptime. + Familiarity with monitoring and continuous-monitoring requirements related to SOC 2, ISO 27001, NIST SP 800-137, or similar frameworks is valuable. + Exposure to OpenTelemetry, eBPF, Sentry, cost-aware telemetry, or cardinality-management techniques is beneficial. + Experience working effectively with AI-assisted development tools and modern agentic engineering workflows is a plus. + Kubernetes experience is useful for supporting the smaller portion of the platform that runs in Kubernetes. + This role is focused on SRE and reliability engineering rather than DevOps ticket management, build-system ownership, cloud cost management, or acting as the on-call team for other engineering squads. ## Description This is a greenfield SRE leadership opportunity within a large-scale, security-focused product environment. You will define what "healthy" means across roughly 70 components spanning cloud services and customer-hosted agents. You will establish SLIs, SLOs, error budgets, monitoring standards, alerting, and escalation practices from the ground up. Your work will directly improve the ability to detect silent security-control degradation before it becomes a widespread customer-impacting issue. You will collaborate closely with engineering leads and senior engineers while building the telemetry and reliability platform yourself. The environment is remote-first, async, technically demanding, and focused on measurable outcomes rather than dashboards for their own sake. This is an opportunity to shape the reliability culture and foundations of a major security product. Accountabilities + Define and establish meaningful SLIs for approximately 70 product components, working with squad leads and senior engineers to agree on ownership, measurement, tiering, SLOs, and error budgets. + Develop a reliability taxonomy covering service availability and latency, fleet reachability and configuration convergence, security-control efficacy, artifact delivery, and telemetry pipeline health. + Ensure reliability indicators are independently measurable and cannot be disabled by the same failure they are intended to detect. + Design and build the telemetry collection pipeline for customer-hosted agents and cloud services, balancing push-based collection, sampling, privacy constraints, data quality, and cardinality. + Extend instrumentation across Python, Go, and Rust components in collaboration with product engineering teams. + Consolidate existing dashboards, queries, and reporting mechanisms into a smaller, more reliable observability platform, retiring tooling that does not provide meaningful operational value. + Implement symptom-based, SLO-driven alerting with multi-window burn-rate principles and clear page, ticket, and dashboard classifications. + Ensure every production alert has a defined owner, documented failure mode, and actionable runbook. + Establish ongoing alert-quality practices, including periodic reviews, measurable actionable-alert rates, and deliberate removal of unnecessary alerts. + Build a machine-readable ownership and escalation model that routes incidents to the appropriate engineering squads. + Establish severity definitions, acknowledgement expectations, follow-the-sun escalation practices, and clean handoff procedures across multiple time zones. + Strengthen incident command and blameless postmortem practices, including reliable timelines, ownership, and follow-through on corrective actions. + Coach engineering squads to own their own operational responsibilities and paging rather than becoming a centralized buffer for other teams' alerts. + Deliver measurable reliability outcomes over the first year, including complete SLI ownership, production telemetry, tiered alerting, squad on-call adoption, and a significant reduction in the time required to detect silent security-control degradation. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [The Road to MLOps: How Verivox Transitioned to AWS](https://www.wearedevelopers.com/videos/1050-the-road-to-mlops-how-verivox-transitioned-to-aws) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [From Doubt to Confidence: How Sentry Uses Verdaccio to Bulletproof SDK Releases](https://www.wearedevelopers.com/videos/739-from-doubt-to-confidence-how-sentry-uses-verdaccio-to-bulletproof-sdk-releases) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [Our GitOps approach for deploying an Identity Provider and an API Gateway in a SaaS company](https://www.wearedevelopers.com/videos/776-our-gitops-approach-for-deploying-an-identity-provider-and-an-api-gateway-in-a-saas-company) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)