Senior Site Reliability Engineer - Storage

Lambda Labs
San Jose, CA, United States
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon S3 Profiling Continuous Integration Data Centers Linux Github General Parallel File Systems Python (Programming Language) Reliability Engineering Ansible
+16 more
Prometheus Transmission Control Protocol (TCP) Ceph (Software) Datadog Data Logging Computer Networking Systems High Performance Computing Grafana Mttr Kubernetes Storage Technologies Sumo Logic (Software) Terraform Docker Jenkins Nvme

Job description

Experteer Overview In this role, you will own the reliability, performance, and capacity of Lambda’s production storage fleet across data centers. You’ll build and maintain monitoring, dashboards, and alerting to track storage health and hardware issues. You’ll investigate incidents using telemetry and performance profiling, and automate triage to focus on root causes. You will design self-healing automation and contribute to CI/CD for storage tooling. This position combines hands-on engineering with cross-team collaboration to scale storage for AI workloads and support Lambda’s mission of ubiquitous compute. Compensation / Benefits * Own reliability, performance, and capacity health of the production storage fleet across data centers * Build and maintain monitoring, dashboards, and alerting for storage performance, capacity, and hardware failures * Investigate and resolve storage incidents using telemetry, logs, and profiling * Automate ticketing, escalation, and incident-response workflows to reduce repetitive triage * Design and maintain self-healing automation for common failure modes (drive replacement, node swaps, rebuild monitoring, capacity rebalancing) * Implement CI/CD pipelines for storage automation and tooling * Collaborate with Storage Engineers, Fleet Orchestration, and Release Engineering to automate deployment/configuration of software-defined storage * Work with hardware and networking teams to diagnose low-level I/O and network issues affecting storage * Participate in on-call rotation to reduce MTTR and avoid repeat paging Tasks * 5+ years operating Linux systems in production or HPC environments * Hands-on storage experience at scale on scale-out or software-defined platforms (e.g., CEPH, Lustre, GPFS) * Hands-on experience with Software-Defined Storage (SDS) platforms and APIs * Strong incident-response instincts with end-to-end ownership of incidents * Experience with monitoring/logging platforms (Prometheus, Grafana, Alertmanager, Datadog, SumoLogic) and dashboards * Experience with Kubernetes and GitOps tooling (ArgoCD, Helm/Kustomize) * Experience with CI/CD tools (GitHub Actions, Jenkins, BuildKite), containers (Docker/Podman), and Python or Go * Experience with Infrastructure as Code (Terraform, Ansible) * Understanding of core storage protocols (NFS, SMB, S3, NVMe-oF/TCP) and related domains Key requirements * generous cash & equity compensation * Health, dental, and vision coverage for you and your dependents * Wellness and commuter stipends * 401k Plan with 2% company match * Flexible paid time off

Requirements

fleet to reduce repetitive triage * Design and maintain self-healing automation for common failure modes (drive replacement, node swaps, rebuild monitoring, capacity rebalancing) * Implement CI/CD pipelines for storage automation and tooling * Collaborate with Storage Engineers, Fleet Orchestration, and Release Engineering to automate deployment/configuration of software-defined storage * Work with hardware and networking teams to diagnose low-level I/O and network issues affecting storage * Participate in on-call rotation to reduce MTTR and avoid repeat paging Tasks * 5+ years operating Linux systems in production or HPC environments * Hands-on storage experience at scale on scale-out or software-defined platforms (e.g., CEPH, Lustre, GPFS) * Hands-on experience with Software-Defined Storage (SDS) platforms and APIs * Strong incident-response instincts with end-to-end ownership of incidents * Experience with monitoring/logging platforms (Prometheus, Grafana, Alertmanager, Datadog, aaaaaaa and and dashboards * Experience with Kubernetes and GitOps tooling (ArgoCD, Helm/Kustomize) * Experience with CI/CD tools (GitHub Actions, Jenkins, BuildKite), containers (Docker/Podman), and Python or Go * Experience with Infrastructure as Code (Terraform, Ansible) * Understanding of core storage protocols (NFS, SMB, S3, NVMe-oF/TCP) and related domains Key requirements * generous cash & equity compensation * Health, dental, and vision coverage for you and your dependents * Wellness and commuter stipends * 401k Plan with 2% company match * Flexible paid time off

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · World Congress 2021

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:07 min

Establishing service level agreements directly for internal platforms

Pawel Piwosz · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all