> Markdown version of [/jobs/ext/2028242-senior-site-reliability-engineer-storage](https://www.wearedevelopers.com/jobs/ext/2028242-senior-site-reliability-engineer-storage). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - Storage - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon S3, Microsoft Azure, Bash Shell, Computer Programming, Distributed File Systems, File Systems, General Parallel File Systems, Information Technology Operations, Python (Programming Language), NetApp Applications, Reliability Engineering, Zabbix, Cloud Platform System, Pure Storage, Storage Technologies, Information Technology, Cloud Integration, Splunk, Golang - **Published:** August 11, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/senior-site-reliability-engineer-storage-santa-clara-ca-usa-58895131 ## About the Role HPC procedures and evaluate distributed file systems * Collaborate with engineering to capture infrastructure requirements and support workflows * Guide methodologies for building, testing, and deploying applications to optimize performance Tasks * BS in Computer Science with 8+ years, MS with 5+ years, or Ph.D. with 3+ years equivalent experience * 8+ years of experience delivering technology solutions and addressing HPC performance bottlenecks * Experience designing, deploying, and managing Enterprise NAS (NetApp, Pure Storage) and S3-based storage (Cloudian MinIO) * Experience with parallel/distributed filesystems (Lustre, GPFS) * Programming/scripting: Python, Bash, Golang * Strong experience with cloud environments (AWS, Azure, or GCP) * Experience with monitoring stacks (Prometheus+Grafana, Elasticsearch+Kibana, Splunk, Zabbix) and emerging tools * Excellent communication and collaboration skills Key requirements * equity * benefits ## Description Experteer Overview As a Senior Site Reliability Engineer, you design, implement, and optimize on-prem HPC storage with cloud integration to support NVIDIA's data-intensive workloads. You'll build scalable storage solutions, automate deployment and management tooling, and ensure reliable IT operations across a growing ecosystem. You will partner with engineering teams to align infrastructure with evolving needs and document best practices for distributed file systems. This role offers a chance to influence storage architectures and enable cutting-edge AI and HPC projects. Join us to help scale performance and drive efficient resource utilization in a world-class, cloud-enabled HPC environment. Compensation / Benefits * Design and implement on-prem HPC infrastructure with cloud extensions * Create scalable storage solutions for data-intensive applications focusing on performance and cost * Develop automation tools for deployment, management, monitoring, and self-service access * Document procedures and evaluate distributed file systems * Collaborate with engineering to capture infrastructure requirements and support workflows * Guide methodologies for building, testing, and deploying applications to optimize performance Tasks * BS in Computer Science with 8+ years, MS with 5+ years, or Ph.D. with 3+ years equivalent experience * 8+ years of experience delivering technology solutions and addressing HPC performance bottlenecks * Experience designing, deploying, and managing Enterprise NAS (NetApp, Pure Storage) and S3-based storage (Cloudian MinIO) * Experience with parallel/distributed filesystems (Lustre, GPFS) * Programming/scripting: Python, Bash, Golang * Strong experience with cloud environments (AWS, Azure, or GCP) * Experience with monitoring stacks (Prometheus+Grafana, Elasticsearch+Kibana, Splunk, Zabbix) and emerging tools * Excellent communication and collaboration skills Key requirements * equity * benefits ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [6 Emerging Technologies We’ll Learn About in 2025](https://www.wearedevelopers.com/magazine/381-6-emerging-technologies-we-ll-learn-about-in-2025)