> Markdown version of [/jobs/ext/3506820-sre-engineer](https://www.wearedevelopers.com/jobs/ext/3506820-sre-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SRE Engineer - **Company:** Número De - **Location:** Vitoria-Gasteiz, Spain - **Salary:** €55,000.0 - €72,000.0 - **Contract:** Permanent contract - **Skills:** Bash Shell, Cloud Computing, Continuous Integration, Data Deduplication, Domain Name System (DNS), Python (Programming Language), Networking Basics, Open Source Technology, Reliability Engineering, Scripting, Load Balancing, Multi-Cloud, Git, Kubernetes, Terraform, Dynatrace - **Published:** September 13, 2026 - **Apply:** https://www.adzuna.es/contact-us.html ## About the Role · 5+ years of experience in SRE, Cloud, or Platform engineering roles · Profound knowledge of SRE principles and incident response practices · Profound knowledge of the observability stack: log aggregation and query design, metrics/dashboard design, and distributed tracing · Profound knowledge of Kubernetes cluster operations and workload objects (Deployments, StatefulSets, Jobs, DaemonSets) and their failure modes · Solid to profound knowledge of networking fundamentals: DNS as infrastructure, TLS certificate lifecycle, load balancing and reverse proxies · Hands-on experience with Infrastructure as Code (Terraform or equivalent) · Profound knowledge of process and OS-level architecture trade-offs as they apply to containers (immutable infrastructure, image strategy) · Scripting proficiency (Bash or Python) for tooling and automation · Strong Git and collaborative workflow experience Nice to Have · Experience with chaos engineering or failure injection programs · Exposure to multi-region or multi-cloud design trade-offs · Familiarity with service mesh implementations · Prior mentoring or technical leadership experience · Certifications such as Site Reliability Engineering (SRE) Foundation or Observability Foundation · Contributions to open-source observability or Kubernetes tooling ## Description We are looking for a Senior SRE Engineer to join our infrastructure team and take technical leadership over production resilience. This role sits at the Senior level on our Cloud/Platform/SRE career path - reliability engineering with a heavy focus on metrics and production systems. You'll define SLIs and SLOs, lead incident response as commander, drive observability strategy end to end, and mentor cloud/platform engineers as you go. You'll work closely with Product and Engineering, balancing speed, quality, and long-term reliability, while making the architectural calls that keep our systems resilient under load. Responsibilities Reliability & Incident Management · Lead incidents as commander: set and revise severity, and know when to mitigate first and diagnose later · Own the incident record and timeline standard, including the link between deployments and incidents · Communicate with stakeholders while an incident is active · Conduct blameless postmortems and drive toil identification and elimination as measured work Observability · Implement the three pillars of observability (logs, metrics, traces) end to end · Design metrics and query strategy - recording rules, dashboard design that separates on-call needs from analyst needs · Define SLIs and SLOs for critical services, choosing the indicator that reflects user experience over the one that's easiest to measure · Design alerting systems - routing, escalation, deduplication, and alert fatigue reduction (multi-window burn-rate alerts) Platform & Production Systems · Design workload health signals - liveness, readiness, and startup probes - and reason about workload lifecycle (SIGTERM handling, termination grace periods, connection draining) · Build runbook automation and self-healing systems to reduce operational toil · Contribute to CI/CD framework improvements and cost optimization initiatives Technical Leadership · Make architectural decisions for reliability-critical systems · Mentor cloud/platform engineers · Influence technical direction on infrastructure and platform decisions ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [JavaScript? No. Java Scripts! - Scripting with Java](https://www.wearedevelopers.com/videos/2094-javascript-no-java-scripts-scripting-with-java) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)