> Markdown version of [/jobs/ext/2025729-senior-site-reliability-engineer-storage](https://www.wearedevelopers.com/jobs/ext/2025729-senior-site-reliability-engineer-storage). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - Storage - **Company:** Lambda Labs - **Location:** San Jose, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon S3, Profiling, Continuous Integration, Data Centers, Linux, Github, General Parallel File Systems, Python (Programming Language), Reliability Engineering, Ansible, Prometheus, Transmission Control Protocol (TCP), Ceph (Software), Datadog, Data Logging, Computer Networking Systems, High Performance Computing, Grafana, Mttr, Kubernetes, Storage Technologies, Sumo Logic (Software), Terraform, Docker, Jenkins, Nvme - **Published:** August 11, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/senior-site-reliability-engineer-storage-san-jose-ca-usa-58891683 ## About the Role fleet to reduce repetitive triage * Design and maintain self-healing automation for common failure modes (drive replacement, node swaps, rebuild monitoring, capacity rebalancing) * Implement CI/CD pipelines for storage automation and tooling * Collaborate with Storage Engineers, Fleet Orchestration, and Release Engineering to automate deployment/configuration of software-defined storage * Work with hardware and networking teams to diagnose low-level I/O and network issues affecting storage * Participate in on-call rotation to reduce MTTR and avoid repeat paging Tasks * 5+ years operating Linux systems in production or HPC environments * Hands-on storage experience at scale on scale-out or software-defined platforms (e.g., CEPH, Lustre, GPFS) * Hands-on experience with Software-Defined Storage (SDS) platforms and APIs * Strong incident-response instincts with end-to-end ownership of incidents * Experience with monitoring/logging platforms (Prometheus, Grafana, Alertmanager, Datadog, aaaaaaa and and dashboards * Experience with Kubernetes and GitOps tooling (ArgoCD, Helm/Kustomize) * Experience with CI/CD tools (GitHub Actions, Jenkins, BuildKite), containers (Docker/Podman), and Python or Go * Experience with Infrastructure as Code (Terraform, Ansible) * Understanding of core storage protocols (NFS, SMB, S3, NVMe-oF/TCP) and related domains Key requirements * generous cash & equity compensation * Health, dental, and vision coverage for you and your dependents * Wellness and commuter stipends * 401k Plan with 2% company match * Flexible paid time off ## Description Experteer Overview In this role, you will own the reliability, performance, and capacity of Lambda's production storage fleet across data centers. You'll build and maintain monitoring, dashboards, and alerting to track storage health and hardware issues. You'll investigate incidents using telemetry and performance profiling, and automate triage to focus on root causes. You will design self-healing automation and contribute to CI/CD for storage tooling. This position combines hands-on engineering with cross-team collaboration to scale storage for AI workloads and support Lambda's mission of ubiquitous compute. Compensation / Benefits * Own reliability, performance, and capacity health of the production storage fleet across data centers * Build and maintain monitoring, dashboards, and alerting for storage performance, capacity, and hardware failures * Investigate and resolve storage incidents using telemetry, logs, and profiling * Automate ticketing, escalation, and incident-response workflows to reduce repetitive triage * Design and maintain self-healing automation for common failure modes (drive replacement, node swaps, rebuild monitoring, capacity rebalancing) * Implement CI/CD pipelines for storage automation and tooling * Collaborate with Storage Engineers, Fleet Orchestration, and Release Engineering to automate deployment/configuration of software-defined storage * Work with hardware and networking teams to diagnose low-level I/O and network issues affecting storage * Participate in on-call rotation to reduce MTTR and avoid repeat paging Tasks * 5+ years operating Linux systems in production or HPC environments * Hands-on storage experience at scale on scale-out or software-defined platforms (e.g., CEPH, Lustre, GPFS) * Hands-on experience with Software-Defined Storage (SDS) platforms and APIs * Strong incident-response instincts with end-to-end ownership of incidents * Experience with monitoring/logging platforms (Prometheus, Grafana, Alertmanager, Datadog, SumoLogic) and dashboards * Experience with Kubernetes and GitOps tooling (ArgoCD, Helm/Kustomize) * Experience with CI/CD tools (GitHub Actions, Jenkins, BuildKite), containers (Docker/Podman), and Python or Go * Experience with Infrastructure as Code (Terraform, Ansible) * Understanding of core storage protocols (NFS, SMB, S3, NVMe-oF/TCP) and related domains Key requirements * generous cash & equity compensation * Health, dental, and vision coverage for you and your dependents * Wellness and commuter stipends * 401k Plan with 2% company match * Flexible paid time off ## Related Videos - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Reliable scalability: How Amazon.com scales on AWS](https://www.wearedevelopers.com/videos/983-reliable-scalability-how-amazon-com-scales-on-aws) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Discover the open source trio you didn’t expect: .NET and PostgreSQL on Linux](https://www.wearedevelopers.com/videos/2042-discover-the-open-source-trio-you-didn-t-expect-net-and-postgresql-on-linux) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [What does the history of data storage tell us about the future?](https://www.wearedevelopers.com/magazine/495-what-does-the-history-of-data-storage-tell-us-about-the-future) - [The Top 10 GitHub Alternatives (2025)](https://www.wearedevelopers.com/magazine/298-the-top-10-github-alternatives-2025)