Site Reliability Engineer (Guardicore Ai Platform) - Remote

Akamai
Valladolid, Spain
4 days ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Bash Shell Databases Continuous Integration Linux DevOps Python (Programming Language) Reliability Engineering Software Tools Cloud Services
+12 more
Prometheus Scripting System Availability Large Language Models Grafana AI Platforms Git Flow Kubernetes Data Pipelines Docker Golang Microservices

Job description

OverviewIn this role you will own the reliability and operational readiness of the Guardicore Data and AI Platform, a cloud-native data/AI platform for security analytics.You will work with cross-functional teams to improve availability, performance, security, and cost efficiency, while guiding engineers on service performance.You will lead complex production investigations and leverage AI-driven automation to streamline operations.This is an opportunity to shape platform reliability at scale within a security-focused, AI-enabled product.”Compensaciones / BeneficiosFlexBase programremote/hybrid/work-from-home optionshealth and well-being benefitsfinancial planning benefitslife beyond work supportResponsabilidadesOperate secure, highly available Kubernetes infrastructure for core microservices, data pipelines, observability, and internal toolingEnhance platform reliability, observability, security, performance, and cost efficiencyProvide guidance to engineers to increase confidence in service performanceLead complex production investigations and drive long-term improvementsLeverage LLMs and AI-driven automation to auto-remediate incidents and streamline operationsPartner across DevOps, Software, Data, AI and Security engineering Teams to investigate and troubleshoot complex problemsParticipate in on-call rotations, guiding restoration and repair of service-impacting issuesRequisitos principales3+ years in SRE, DevOps, or Platform Engineering with proven troubleshooting of complex systemsDesign and implement monitoring/observability strategy using Prometheus and GrafanaProduction experience with Kubernetes, Docker, Helm, and cloud providers (GCP, Azure, Linode, AWS) on LinuxExceptional troubleshooting across network, system, applications, and databasesExperience with GitOps, CI/CD, and Infrastructure as CodeScripting/programming in Python, Go, and BashExperience using AI tools in operations and proposing automation initiativesTechnical leadership and ownership in cross-team initiativestechnical leadershipownershipcross-functional collaborationKubernetesDockerHelm

Requirements

Requisitos principales3+ years in SRE, DevOps, or Platform Engineering with proven troubleshooting of complex systems Design and implement monitoring/observability strategy using Prometheus and Grafana Production experience with Kubernetes, Docker, Helm, and cloud providers (GCP, Azure, Linode, AWS) on Linux Exceptional troubleshooting across network, system, applications, and databases Experience with GitOps, CI/CD, and Infrastructure as Code Scripting/programming in Python, Go, and Bash Experience using AI tools in operations and proposing automation initiatives Technical leadership and ownership in cross-team initiatives technical leadership ownership cross-functional collaboration Kubernetes Docker

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:06 min

Empowering site reliability engineers with integrated AI agents

Osmar Matos Osmar Matos · World Congress 2026 Europe

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all