Senior Site Reliability Engineer - Storage
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+8 more
Job description
Experteer Overview As a Senior Site Reliability Engineer, you design, implement, and optimize on-prem HPC storage with cloud integration to support NVIDIA’s data-intensive workloads. You’ll build scalable storage solutions, automate deployment and management tooling, and ensure reliable IT operations across a growing ecosystem. You will partner with engineering teams to align infrastructure with evolving needs and document best practices for distributed file systems. This role offers a chance to influence storage architectures and enable cutting-edge AI and HPC projects. Join us to help scale performance and drive efficient resource utilization in a world-class, cloud-enabled HPC environment. Compensation / Benefits * Design and implement on-prem HPC infrastructure with cloud extensions * Create scalable storage solutions for data-intensive applications focusing on performance and cost * Develop automation tools for deployment, management, monitoring, and self-service access * Document procedures and evaluate distributed file systems * Collaborate with engineering to capture infrastructure requirements and support workflows * Guide methodologies for building, testing, and deploying applications to optimize performance Tasks * BS in Computer Science with 8+ years, MS with 5+ years, or Ph.D. with 3+ years equivalent experience * 8+ years of experience delivering technology solutions and addressing HPC performance bottlenecks * Experience designing, deploying, and managing Enterprise NAS (NetApp, Pure Storage) and S3-based storage (Cloudian MinIO) * Experience with parallel/distributed filesystems (Lustre, GPFS) * Programming/scripting: Python, Bash, Golang * Strong experience with cloud environments (AWS, Azure, or GCP) * Experience with monitoring stacks (Prometheus+Grafana, Elasticsearch+Kibana, Splunk, Zabbix) and emerging tools * Excellent communication and collaboration skills Key requirements * equity * benefits
Requirements
HPC procedures and evaluate distributed file systems * Collaborate with engineering to capture infrastructure requirements and support workflows * Guide methodologies for building, testing, and deploying applications to optimize performance Tasks * BS in Computer Science with 8+ years, MS with 5+ years, or Ph.D. with 3+ years equivalent experience * 8+ years of experience delivering technology solutions and addressing HPC performance bottlenecks * Experience designing, deploying, and managing Enterprise NAS (NetApp, Pure Storage) and S3-based storage (Cloudian MinIO) * Experience with parallel/distributed filesystems (Lustre, GPFS) * Programming/scripting: Python, Bash, Golang * Strong experience with cloud environments (AWS, Azure, or GCP) * Experience with monitoring stacks (Prometheus+Grafana, Elasticsearch+Kibana, Splunk, Zabbix) and emerging tools * Excellent communication and collaboration skills Key requirements * equity * benefits
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
7 Cloud Computing Trends Coming in 2025 for Developers
Dev Digest 120 - Apple and peers
What Are The Top Skills Required For Azure Developers?
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence