Senior Site Reliability Engineer - Storage
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+16 more
Job description
Experteer Overview In this role, you will own the reliability, performance, and capacity of Lambda’s production storage fleet across data centers. You’ll build and maintain monitoring, dashboards, and alerting to track storage health and hardware issues. You’ll investigate incidents using telemetry and performance profiling, and automate triage to focus on root causes. You will design self-healing automation and contribute to CI/CD for storage tooling. This position combines hands-on engineering with cross-team collaboration to scale storage for AI workloads and support Lambda’s mission of ubiquitous compute. Compensation / Benefits * Own reliability, performance, and capacity health of the production storage fleet across data centers * Build and maintain monitoring, dashboards, and alerting for storage performance, capacity, and hardware failures * Investigate and resolve storage incidents using telemetry, logs, and profiling * Automate ticketing, escalation, and incident-response workflows to reduce repetitive triage * Design and maintain self-healing automation for common failure modes (drive replacement, node swaps, rebuild monitoring, capacity rebalancing) * Implement CI/CD pipelines for storage automation and tooling * Collaborate with Storage Engineers, Fleet Orchestration, and Release Engineering to automate deployment/configuration of software-defined storage * Work with hardware and networking teams to diagnose low-level I/O and network issues affecting storage * Participate in on-call rotation to reduce MTTR and avoid repeat paging Tasks * 5+ years operating Linux systems in production or HPC environments * Hands-on storage experience at scale on scale-out or software-defined platforms (e.g., CEPH, Lustre, GPFS) * Hands-on experience with Software-Defined Storage (SDS) platforms and APIs * Strong incident-response instincts with end-to-end ownership of incidents * Experience with monitoring/logging platforms (Prometheus, Grafana, Alertmanager, Datadog, SumoLogic) and dashboards * Experience with Kubernetes and GitOps tooling (ArgoCD, Helm/Kustomize) * Experience with CI/CD tools (GitHub Actions, Jenkins, BuildKite), containers (Docker/Podman), and Python or Go * Experience with Infrastructure as Code (Terraform, Ansible) * Understanding of core storage protocols (NFS, SMB, S3, NVMe-oF/TCP) and related domains Key requirements * generous cash & equity compensation * Health, dental, and vision coverage for you and your dependents * Wellness and commuter stipends * 401k Plan with 2% company match * Flexible paid time off
Requirements
fleet to reduce repetitive triage * Design and maintain self-healing automation for common failure modes (drive replacement, node swaps, rebuild monitoring, capacity rebalancing) * Implement CI/CD pipelines for storage automation and tooling * Collaborate with Storage Engineers, Fleet Orchestration, and Release Engineering to automate deployment/configuration of software-defined storage * Work with hardware and networking teams to diagnose low-level I/O and network issues affecting storage * Participate in on-call rotation to reduce MTTR and avoid repeat paging Tasks * 5+ years operating Linux systems in production or HPC environments * Hands-on storage experience at scale on scale-out or software-defined platforms (e.g., CEPH, Lustre, GPFS) * Hands-on experience with Software-Defined Storage (SDS) platforms and APIs * Strong incident-response instincts with end-to-end ownership of incidents * Experience with monitoring/logging platforms (Prometheus, Grafana, Alertmanager, Datadog, aaaaaaa and and dashboards * Experience with Kubernetes and GitOps tooling (ArgoCD, Helm/Kustomize) * Experience with CI/CD tools (GitHub Actions, Jenkins, BuildKite), containers (Docker/Podman), and Python or Go * Experience with Infrastructure as Code (Terraform, Ansible) * Understanding of core storage protocols (NFS, SMB, S3, NVMe-oF/TCP) and related domains Key requirements * generous cash & equity compensation * Health, dental, and vision coverage for you and your dependents * Wellness and commuter stipends * 401k Plan with 2% company match * Flexible paid time off
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Dev Digest 120 - Apple and peers
Learning Kubernetes made easy with KubeCampus
Dev Digest 121 - AI goes offline
What does the history of data storage tell us about the future?