Site Reliability Engineer

Mlabs Ltd
Madrid, Spain
22 days ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Compensation
€150,000.0 - €200,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Systems Engineering Microsoft Azure Computer Programming Continuous Integration Distributed Systems Python (Programming Language) Reliability Engineering Azure Active Directory Prometheus Systems Integration Cloud Platform System
+4 more
Grafana Kubernetes Deployment Automation Terraform

Job description

Our client is seeking a Senior Site Reliability Engineer (Azure) to architect and scale a robust infrastructure foundation for a high-growth distributed systems platform. This position is critical for ensuring that the platform operates as a secure, scalable, and production-ready environment capable of supporting complex enterprise use cases and high reliability standards. The successful candidate will take a lead role in designing infrastructure from first principles, bridging the gap between product requirements and technical execution. This is a high-impact opportunity for a seasoned engineer to build greenfield Azure environments and establish operational excellence across a global ecosystem., + Infrastructure Design: Architect and deploy secure, scalable Azure infrastructure tailored for production-grade distributed systems.

  • Automation & IaC: Develop and maintain Terraform-based infrastructure as code to enable repeatable, automated deployments across various environments.
  • Technical Leadership: Translate ambiguous product and customer requirements into structured technical architecture and actionable execution plans.
  • Platform Enhancement: Build and optimize platform services, APIs, and integrations to extend core system capabilities.
  • Cross-Functional Collaboration: Partner with engineering, security, and product teams to deliver enterprise-ready infrastructure solutions.
  • Operational Excellence: Drive improvements in reliability, observability, and incident response while providing Tier 2 infrastructure support for customer deployments. Performance Goals: What Success Looks Like In the first 6-12 months of this role, the following milestones are expected to be achieved:

  • Production Readiness: The Azure environment is established as a fully production-ready deployment setting for the platform.
  • Scalable Deployments: All customer deployments are verified as repeatable, scalable, and secure.
  • Feature Parity: Azure achieves full feature parity with all other supported cloud environments within the organization’s ecosystem. Interview Process
  1. Recruiter & Technical Screening: Initial HR call followed by an introductory technical interview covering foundational questions.
  2. Hiring Manager Interview: A deeper dive into experience, alignment, and role-specific expectations.
  3. Technical Interview: A comprehensive evaluation of architectural and technical execution skills.
  4. Final Leadership Interview: A concluding session with the VP of Engineering.

Requirements

  • Proven Track Record: Extensive experience designing and building production-grade systems specifically on the Azure stack.
  • Problem Solving: Ability to transform high-level requirements into scalable, delivered systems.
  • Communication: Strong technical communication skills with the ability to interface with both engineering teams and non-technical stakeholders.
  • Mindset: A high-ownership approach with a strong bias for action and accountability. Functional Expertise

  • Azure Services: Deep knowledge of Azure networking, compute, identity, security, and storage.
  • Infrastructure as Code: Advanced proficiency with Terraform at production scale.
  • Programming: Professional experience in Go and/or Python.
  • Systems Engineering: Background in distributed systems, high-availability architectures, or platform engineering.
  • CI/CD: Experience with automation tooling for the entire infrastructure lifecycle. Preferred Qualifications

  • Hands-on experience with Kubernetes and container orchestration.
  • Familiarity with observability tools such as Prometheus and Grafana.
  • Experience with workflow/orchestration platforms like Argo or Spacelift.

Benefits & conditions

Our client offers a competitive compensation package designed to reward high-impact contributors, including:

  • Equity & Tokens: Participation in the long-term growth of the project.
  • Performance Bonuses: Annual incentives based on individual and company milestones.
  • Health & Retirement: Comprehensive health insurance and 401k plans (available for US-based employees).

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:05 min

Measuring system availability utilizing Prometheus and straightforward PromQL

Alexander Schwartz Alexander Schwartz · WWC 2025

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

13:07 min

Configuring application observability with Micrometer and Prometheus

Aleksandr Kalikov · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · WWC 2025

Videos

See all

Related articles

See all