Site Reliability Engineer

Kyndryl
Madrid, Spain
4 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
10 years minimum
Working hours
Regular working hours

Tech stack

Microsoft Windows IBM AIX Amazon Web Services Microsoft Azure Bash Shell Unix Information Systems Linux HP Systems Insight Manager JSON Python (Programming Language) OpenShift
+14 more
Windows PowerShell Red Hat Enterprise Linux Reliability Engineering Ansible Prometheus Solaris (Operating System) YAML Scripting Google Cloud Grafana Kubernetes Information Technology Terraform Devsecops

Job description

Experteer Overview As a Site Reliability Engineer at Kyndryl, you will drive reliability, resiliency, and continuous improvement across enterprise systems.You’ll work with a cross?functional team to analyze business needs, design robust solutions, and guide end?to?end services that span customer sites and platforms.You’ll build and maintain monitoring, manage incidents, and implement innovative tools to enhance performance and security.This is a chance to shape scalable, trusted services within a collaborative, growth?minded culture.Compensaciones / Beneficios - Drive reliability and resiliency across information systems and ecosystems - Design and implement monitoring to meet reliability and performance targets - Lead incident management and escalate where needed - Define service level indicators and objectives with stakeholders - Develop and deploy end?to?end services spanning customer sites and platforms - Introduce tools and practices to improve operations and gather feedback - Collaborate with DevSecOps, development, and operations teams - Proactively identify and mitigate operational issues - Foster customer relationships to enable trusted partnerships - Support scalable, secure, and robust service deliveryResponsabilidades - 10+ years of experience in operational management including incident management and escalations - Experience designing and implementing application monitoring for reliability and performance - Experience implementing strategies to cap operations load and handle overflow with appropriate tooling and metrics - Enterprise environment experience: Windows, Linux (RHEL preferred), UNIX (AIX, Solaris); storage and hyperscalercloud (AWS, Azure, GCP) - Experience with data formats and scripting languages JSON, YAML, Bash and/or PowerShell - BS degree in Computer Science, Engineering, or related technical field (preferred) - Expertise with Ansible, Terraform, and Python - Experience with distributed technologies and Kubernetes - Experience with open?source tooling such as Prometheus, Grafana, Loki - Public cloud platforms and OpenShift knowledgeRequisitos principales - hybrid?friendly culture - Be Well programs for financial, mental, physical, and social health - opportunities for certifications and learning (Microsoft, Google, Amazon) - personalized career development and coaching - continuous feedback culture - global collaboration opportunities

Requirements

This is a chance to shape scalable, trusted services within a collaborative, growth?minded culture.Compensaciones / Beneficios - Drive reliability and resiliency across information systems and ecosystems - Design and implement monitoring to meet reliability and performance targets - Lead incident management and escalate where needed - Define service level indicators and objectives with stakeholders - Develop and deploy end?to?end services spanning customer sites and platforms - Introduce tools and practices to improve operations and gather feedback - Collaborate with DevSecOps, development, and operations teams - Proactively identify and mitigate operational issues - Foster customer relationships to enable trusted partnerships - Support scalable, secure, and robust service deliveryResponsabilidades - 10+ years of experience in operational management including incident management and escalations - Experience designing and implementing application monitoring for reliability and performance - Experience implementing strategies to cap operations load and handle overflow with appropriate tooling and metrics - Enterprise environment experience: Windows, Linux (RHEL preferred), UNIX (AIX, Solaris); storage and hyperscalercloud (AWS, Azure, GCP) - Experience with data formats and scripting languages JSON, YAML, Bash and/or PowerShell - BS degree in Computer Science, Engineering, or related technical field (preferred) - Expertise with Ansible, Terraform, and Python - Experience with distributed technologies and Kubernetes - Experience with open?source tooling such as Prometheus, Grafana, Loki - Public cloud platforms and OpenShift knowledgeRequisitos principales - hybrid?friendly culture - Be Well programs for financial, mental, physical, and social health - opportunities for certifications and learning (Microsoft, Google, Amazon) - personalized career development and coaching - continuous feedback culture - global collaboration opportunities

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:03 min

Microsoft integrating native Unix coreutils into Windows environments

Chris Heilmann +2 · LIVE

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

1:38 min

Managing and versioning system prompts as YAML files

Kevin Lewis Kevin Lewis +1 · WWC 2025

1:06 min

Developer experience and project variety at scale

Alexandra Petri · WWC 2023

2:04 min

Defining timestamps and the international standard format

Denny Biasiolli Denny Biasiolli · Europe 2026 Virtual

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

Videos

See all

Related articles

See all