> Markdown version of [/jobs/ext/3105594-site-reliability-engineer-guardicore-ai-platform-remote](https://www.wearedevelopers.com/jobs/ext/3105594-site-reliability-engineer-guardicore-ai-platform-remote). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer (Guardicore Ai Platform) - Remote - **Company:** Akamai - **Location:** Lleida (Lérida), Spain (Remote available) - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Bash Shell, Databases, Continuous Integration, Linux, DevOps, Python (Programming Language), Reliability Engineering, Software Tools, Cloud Services, Prometheus, Scripting, System Availability, Large Language Models, Grafana, AI Platforms, Git Flow, Kubernetes, Data Pipelines, Docker, Golang, Microservices - **Published:** September 27, 2026 - **Apply:** https://www.buscojobs.com.es/site-reliability-engineer-guardicore-ai-platform-remote-en-lleida-ID-372566880 ## About the Role Requisitos principales 3+ years in SRE, DevOps, or Platform Engineering with proven troubleshooting of complex systems Design and implement monitoring/observability strategy using Prometheus and Grafana Production experience with Kubernetes, Docker, Helm, and cloud providers (GCP, Azure, Linode, AWS) on Linux Exceptional troubleshooting across network, system, applications, and databases Experience with GitOps, CI/CD, and Infrastructure as Code Scripting/programming in Python, Go, and Bash Experience using AI tools in operations and proposing automation initiatives Technical leadership and ownership in cross-team initiatives technical leadership ownership cross-functional collaboration Kubernetes Docker ## Description OverviewIn this role you will own the reliability and operational readiness of the Guardicore Data and AI Platform, a cloud-native data/AI platform for security analytics. You will work with cross-functional teams to improve availability, performance, security, and cost efficiency, while guiding engineers on service performance. You will lead complex production investigations and leverage AI-driven automation to streamline operations. This is an opportunity to shape platform reliability at scale within a security-focused, AI-enabled product."Compensaciones / Beneficios FlexBase programremote/hybrid/work-from-home optionshealth and well-being benefitsfinancial planning benefitslife beyond work supportResponsabilidades Operate secure, highly available Kubernetes infrastructure for core microservices, data pipelines, observability, and internal toolingEnhance platform reliability, observability, security, performance, and cost efficiencyProvide guidance to engineers to increase confidence in service performanceLead complex production investigations and drive long-term improvementsLeverage LLMs and AI-driven automation to auto-remediate incidents and streamline operationsPartner across DevOps, Software, Data, AI and Security engineering Teams to investigate and troubleshoot complex problemsParticipate in on-call rotations, guiding restoration and repair of service-impacting issuesRequisitos principales 3+ years in SRE, DevOps, or Platform Engineering with proven troubleshooting of complex systemsDesign and implement monitoring/observability strategy using Prometheus and GrafanaProduction experience with Kubernetes, Docker, Helm, and cloud providers (GCP, Azure, Linode, AWS) on LinuxExceptional troubleshooting across network, system, applications, and databasesExperience with GitOps, CI/CD, and Infrastructure as CodeScripting/programming in Python, Go, and BashExperience using AI tools in operations and proposing automation initiativesTechnical leadership and ownership in cross-team initiativestechnical leadershipownershipcross-functional collaborationKubernetesDockerHelm ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)