Site Reliability Engineer (Guardicore Ai Platform) - Remote
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+12 more
Job description
OverviewIn this role you will own the reliability and operational readiness of the Guardicore Data and AI Platform, a cloud-native data/AI platform for security analytics.You will work with cross-functional teams to improve availability, performance, security, and cost efficiency, while guiding engineers on service performance.You will lead complex production investigations and leverage AI-driven automation to streamline operations.This is an opportunity to shape platform reliability at scale within a security-focused, AI-enabled product.”Compensaciones / BeneficiosFlexBase programremote/hybrid/work-from-home optionshealth and well-being benefitsfinancial planning benefitslife beyond work supportResponsabilidadesOperate secure, highly available Kubernetes infrastructure for core microservices, data pipelines, observability, and internal toolingEnhance platform reliability, observability, security, performance, and cost efficiencyProvide guidance to engineers to increase confidence in service performanceLead complex production investigations and drive long-term improvementsLeverage LLMs and AI-driven automation to auto-remediate incidents and streamline operationsPartner across DevOps, Software, Data, AI and Security engineering Teams to investigate and troubleshoot complex problemsParticipate in on-call rotations, guiding restoration and repair of service-impacting issuesRequisitos principales3+ years in SRE, DevOps, or Platform Engineering with proven troubleshooting of complex systemsDesign and implement monitoring/observability strategy using Prometheus and GrafanaProduction experience with Kubernetes, Docker, Helm, and cloud providers (GCP, Azure, Linode, AWS) on LinuxExceptional troubleshooting across network, system, applications, and databasesExperience with GitOps, CI/CD, and Infrastructure as CodeScripting/programming in Python, Go, and BashExperience using AI tools in operations and proposing automation initiativesTechnical leadership and ownership in cross-team initiativestechnical leadershipownershipcross-functional collaborationKubernetesDockerHelm
Requirements
Requisitos principales3+ years in SRE, DevOps, or Platform Engineering with proven troubleshooting of complex systems Design and implement monitoring/observability strategy using Prometheus and Grafana Production experience with Kubernetes, Docker, Helm, and cloud providers (GCP, Azure, Linode, AWS) on Linux Exceptional troubleshooting across network, system, applications, and databases Experience with GitOps, CI/CD, and Infrastructure as Code Scripting/programming in Python, Go, and Bash Experience using AI tools in operations and proposing automation initiatives Technical leadership and ownership in cross-team initiatives technical leadership ownership cross-functional collaboration Kubernetes Docker
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Find a Developer Job: 12 Best Job Sites For Developers
Dev Digest 120 - Apple and peers
Dev Digest 121 - AI goes offline
Is Software Engineering Over-Saturated?