Senior Site Reliability Engineer - Storage

NVIDIA Ltd.
Santa Clara, CA, United States
2 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Amazon Web Services Amazon S3 Microsoft Azure Bash Shell Computer Programming Distributed File Systems File Systems General Parallel File Systems Information Technology Operations Python (Programming Language) NetApp Applications Reliability Engineering
+8 more
Zabbix Cloud Platform System Pure Storage Storage Technologies Information Technology Cloud Integration Splunk Golang

Job description

Experteer Overview As a Senior Site Reliability Engineer, you design, implement, and optimize on-prem HPC storage with cloud integration to support NVIDIA’s data-intensive workloads. You’ll build scalable storage solutions, automate deployment and management tooling, and ensure reliable IT operations across a growing ecosystem. You will partner with engineering teams to align infrastructure with evolving needs and document best practices for distributed file systems. This role offers a chance to influence storage architectures and enable cutting-edge AI and HPC projects. Join us to help scale performance and drive efficient resource utilization in a world-class, cloud-enabled HPC environment. Compensation / Benefits * Design and implement on-prem HPC infrastructure with cloud extensions * Create scalable storage solutions for data-intensive applications focusing on performance and cost * Develop automation tools for deployment, management, monitoring, and self-service access * Document procedures and evaluate distributed file systems * Collaborate with engineering to capture infrastructure requirements and support workflows * Guide methodologies for building, testing, and deploying applications to optimize performance Tasks * BS in Computer Science with 8+ years, MS with 5+ years, or Ph.D. with 3+ years equivalent experience * 8+ years of experience delivering technology solutions and addressing HPC performance bottlenecks * Experience designing, deploying, and managing Enterprise NAS (NetApp, Pure Storage) and S3-based storage (Cloudian MinIO) * Experience with parallel/distributed filesystems (Lustre, GPFS) * Programming/scripting: Python, Bash, Golang * Strong experience with cloud environments (AWS, Azure, or GCP) * Experience with monitoring stacks (Prometheus+Grafana, Elasticsearch+Kibana, Splunk, Zabbix) and emerging tools * Excellent communication and collaboration skills Key requirements * equity * benefits

Requirements

HPC procedures and evaluate distributed file systems * Collaborate with engineering to capture infrastructure requirements and support workflows * Guide methodologies for building, testing, and deploying applications to optimize performance Tasks * BS in Computer Science with 8+ years, MS with 5+ years, or Ph.D. with 3+ years equivalent experience * 8+ years of experience delivering technology solutions and addressing HPC performance bottlenecks * Experience designing, deploying, and managing Enterprise NAS (NetApp, Pure Storage) and S3-based storage (Cloudian MinIO) * Experience with parallel/distributed filesystems (Lustre, GPFS) * Programming/scripting: Python, Bash, Golang * Strong experience with cloud environments (AWS, Azure, or GCP) * Experience with monitoring stacks (Prometheus+Grafana, Elasticsearch+Kibana, Splunk, Zabbix) and emerging tools * Excellent communication and collaboration skills Key requirements * equity * benefits

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

3:37 min

Why differing legacy workflows complicate monitoring tool migrations

Mathias Palmersheim Mathias Palmersheim · Europe 2026 Virtual

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

Videos

See all

Related articles

See all