Site Reliability Engineer

Ai-driven
Ventas, Spain
1 day ago
Apply on www.adzuna.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Systems Engineering Bash Shell Software Debugging Linux DevOps Distributed Systems Github Reliability Engineering Runbook
+12 more
Software Engineering Systems Architecture Data Logging Scripting System Availability Delivery Pipeline Gitlab Kubernetes Infrastructure Automation Frameworks Deployment Automation Terraform Serverless Computing

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer (Remote Build) based in Spain. This role sits at the core of a next-generation platform enabling AI-driven services across global employment infrastructure. You will be responsible for ensuring the reliability, scalability, and security of systems that power complex integrations across multiple countries and compliance frameworks. The environment is highly distributed, async-first, and built for engineers who thrive in ownership and autonomy. You will design and operate infrastructure that supports Kubernetes-based deployments, cloud-native services, and agentic workflows at scale. This is a hands-on engineering role where you will directly influence system architecture, operational practices, and developer experience. You will collaborate closely with engineering, product, and security teams to ensure performance, cost efficiency, and operational excellence across all services. Accountabilities: You will design and maintain scalable infrastructure-as-code solutions using tools like Terraform and Kubernetes, ensuring robust, repeatable, and secure deployments across environments. You will also support platform evolution by improving automation and deployment workflows. You will build and operate observability systems including monitoring, logging, and alerting, while leading incident response, postmortems, and reliability improvements to ensure high system availability. You will embed security and compliance practices into infrastructure and operational workflows, ensuring adherence to global regulatory requirements while minimizing friction for engineering teams. You will optimize system performance, reliability, and cloud costs through continuous analysis and tuning of infrastructure and workloads across distributed systems. You will eliminate operational toil by developing automation tools and scalable processes that reduce manual intervention and improve engineering efficiency. You will partner with product and platform teams to improve APIs, deployment systems, and developer experience, ensuring infrastructure supports long-term scalability and maintainability.

Requirements

You bring senior-level experience in Site Reliability Engineering, DevOps, or Systems Engineering, with a proven track record of operating production systems at scale in cloud environments. You are comfortable owning reliability end-to-end. You have deep hands-on expertise with Kubernetes and AWS, including networking, compute, storage, and managed services, and understand how to operate resilient distributed systems. You are highly proficient with infrastructure-as-code tools such as Terraform and apply software engineering principles to infrastructure design and management. You have strong experience with CI/CD pipelines and deployment automation using tools like GitHub Actions, GitLab, or similar, including rollback strategies and safe deployment practices. You are comfortable working with Linux systems, debugging production issues, writing scripts (especially in Bash), and understanding system-level behavior. You are an effective communicator who can translate complex infrastructure concepts into clear explanations, documentation, and runbooks for both technical and non-technical stakeholders.

Benefits & conditions

Competitive salary aligned with global benchmarks and experience level Fully remote work with flexible scheduling and async-first culture Equity or stock option opportunities depending on role eligibility Flexible paid time off and generous parental leave policies Learning and development budget to support continuous growth Home office and equipment support to set up your workspace Mental health and wellness support services Opportunities to work on globally distributed, high-impact infrastructure systems How Jobgether works

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · World Congress 2021

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

1:06 min

Empowering site reliability engineers with integrated AI agents

Osmar Matos Osmar Matos · World Congress 2026 Europe

Videos

See all

Related articles

See all