> Markdown version of [/jobs/ext/1680070-site-reliability-engineer-intermediate-to-senior-staff](https://www.wearedevelopers.com/jobs/ext/1680070-site-reliability-engineer-intermediate-to-senior-staff). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer Intermediate to Senior Staff - **Company:** GitLab - **Location:** Madrid, Spain - **Salary:** €126,400.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Build Automation, Continuous Integration, Software Debugging, Ruby, Software Engineering, Data Logging, Gitlab, Kubernetes, Terraform - **Published:** July 13, 2026 - **Apply:** https://www.jobleads.com/es/job/e92732cfaf80be2c9a4ed9a1a48e041c0 ## About the Role * Experience keeping production systems reliable, combining an operations mindset with real software engineering practice. * Experience building net-new infrastructure tooling and automation, not just configuring existing tools (e.g., Terraform modules, Kubernetes operators or controllers, or production automation and services written from scratch). * The ability to read, debug, and reason about code. Most of our teams work in Go; some work in Ruby. You can discuss a piece of code's behavior, performance, and failure modes. * Experience with infrastructure as code, and with Kubernetes and its ecosystem, at a depth appropriate to your level. * Hands-on experience with at least one major cloud provider (GCP or AWS). * Familiarity with observability practices, including metrics, logging, alerting, and SLOs or SLIs, and using data to inform operational decisions. * Comfort participating in on-call and incident response, with a structured approach to troubleshooting under pressure. * Strong written communication and the ability to operate as a manager-of-one in an async, distributed environment. * A track record of using automation, and increasingly AI, to reduce toil and improve how you and your team work. * Alignment with GitLab's values and a commitment to working in accordance with them. ## Description + You make meaningful contributions to reliability, automation, and operational efficiency, working independently within a scoped area. + You diagnose issues on your own, understand system dependencies, and can explain the tradeoffs you made. + You prioritize well, break work into manageable steps, and use automation to reduce toil. + You document your work clearly and keep yourself moving without needing check-ins. * Senior + You drive reliability improvements across multiple projects or services and prioritize them based on real system needs. + You lead investigations, anticipate cascading failures, and coordinate incident response. + You own delivery end to end, unblock others, and improve the patterns your team works by. + You communicate complex ideas clearly, influence how work gets done, and enable coordination across teams. * Staff + You shape reliability strategy across teams and services and define patterns that others reuse. + You introduce prevention strategies, identify systemic weaknesses, and influence incident response practices beyond your immediate area. + You design execution and automation approaches that work at organizational scale. + You connect reliability work to platform and business needs. * Senior Staff + You set technical direction for reliability across a sub-department, not just a team. + You drive the hardest, most ambiguous systems problems and establish standards and guardrails that multiple teams adopt. + You mentor Staff and Senior engineers. + You align reliability strategy with long-range platform direction and represent Infrastructure's interests across the wider Engineering organization. What you'll do * Keep user-facing services and production systems reliable, scalable, and efficient. * Build automation and tooling that reduces toil and replaces manual work with repeatable, infrastructure-as-code-driven workflows. * Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling. * Write and maintain infrastructure as code, and ship changes safely through CI/CD and GitOps. * Participate in on-call, triage alerts, follow and improve runbooks, and elevate appropriately. * Contribute to the observability stack, using metrics, logs, and SLOs to detect symptoms early rather than just outages. * Take part in incident response and post-incident reviews, turning learnings into changes in automation and process. * Document runbooks, architecture decisions, and reviews so your findings become repeatable practices., Infrastructure Platforms is responsible for the availability, reliability, performance, and scalability of GitLab's user-facing services, most notably GitLab.com. The department spans sub-departments including Production Engineering and Dedicated, and the teams within them own everything from the production fleet and networking platform to observability, incident response, and our single-tenant Dedicated offering. We are a globally distributed, all-remote group that works asynchronously, favors automation over toil, and closes the loop with monitoring and metrics to drive accountability. ## Related Videos - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [GitOps for the people](https://www.wearedevelopers.com/videos/815-gitops-for-the-people) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) ## Related Articles - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [The Top 10 GitHub Alternatives (2025)](https://www.wearedevelopers.com/magazine/298-the-top-10-github-alternatives-2025) - [Stop Googling Git Commands. Start Actually Learning Git.](https://www.wearedevelopers.com/magazine/730-stop-googling-git-commands-start-actually-learning-git) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)