Senior Storage Software Engineer - DGX Cloud

NVIDIA Corporation
Santa Clara, CA, United States
3 days ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$224,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon S3 C++ (Programming Language) Cloud Computing Computer Clusters Software Debugging Distributed File Systems File Systems InfiniBand Python (Programming Language) Linux Kernel Metadata
+11 more
Open Source Technology Performance Tuning Remote Direct Memory Access Graphics Processing Unit (GPU) Kubernetes Storage Technologies Information Technology Production Code Hardware Infrastructure Block Storage Nvme

Job description

NVIDIA DGXC Storage team handles some of the fastest training and inference tasks. Every GPU cycle depends on a storage platform built to keep tens of thousands of accelerators continuously busy. It maintains exabytes of data securely and powers the largest AI workloads worldwide across cloud, neocloud, and on-prem setups. With the growth of accelerated computing, storage is essential. It can make the difference between effective GPU use and wasted potential, and between launching a frontier model on time or missing the deadline by months. We’re looking for a hands-on Storage Software Engineer to join the storage team as an individual contributor and technical lead. You will contribute to open-source parallel and distributed file systems and keep our largest GPU clusters fast, reliable, and durable. You will stay deeply hands-on: writing and reviewing production code, chasing root causes in the field, and setting the configuration and tuning standards our GPU fleets run on. This is a chance to do foundational storage engineering for the AI era at the company that introduced accelerated computing.

What you’ll be doing:

  • Contribute to open-source file systems. Contribute code to open-source parallel and distributed file systems, and distributed object storage. Upstream fixes and features, and engage directly with the upstream communities and maintainers.
  • Serve as a hands-on storage software lead. Write and review production code yourself, and read kernel, NFS, NVMe-oF, or SPDK source when a bug requires it. Make the final technical calls on storage deliveries against measurable targets.
  • Triage and troubleshoot at scale. Triage, troubleshoot, and root-cause large, complex storage issues across very large GPU clusters (tens of thousands of GPUs) - I/O and metadata performance, data corruption, and recovery.
  • Validate architecture and capabilities. Validate storage architecture, capabilities, performance, and durability. Run scale tests, benchmarks, and recovery drills, and qualify new builds against measurable performance and durability targets.
  • Recommend configuration, tuning, and guidelines. Define and recommend configuration, tuning, and operational best practices for high-performance file systems on GPU infrastructure, and help operators and internal customers apply them.
  • Partner broadly. Work with training, inference, and accelerated-computing teams, site-reliability and operations, networking, and security, and collaborate with cloud providers, neocloud operators, and storage vendors on a common architecture.
  • Work AI-first. Use modern AI coding and agentic tools day-to-day to accelerate building, debugging, validation, and operations.

Requirements

  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field - or equivalent experience. Over 12 years of direct experience in storage software engineering, including extensive involvement with a high-performance parallel or distributed file system handling multi-petabyte scale.
  • Contributions to open-source projects involving a distributed or parallel file system. You are fully engaged in engineering tasks. You write and review production code, examine file system, kernel, NVMe-oF, or SPDK source to identify bugs, and personally conduct scale tests or recovery drills instead of assigning them to others.
  • Experience diagnosing and resolving storage problems in extensive GPU or HPC clusters, including analysis of I/O and metadata performance.
  • Strong proficiency in at least one systems language (C, C++, Rust, or Go) and proficiency in Python; comfortable in the Linux kernel storage and networking stacks (block layer, RDMA / RoCE / InfiniBand, NVMe, page cache, VFS, multipath).
  • Solid understanding of object storage (S3 / Swift-class) and block storage (NVMe-oF, iSCSI).
  • Strong written and verbal communication; capable of clarifying complex technical trade-offs to engineers, SREs, vendors, and internal customers.
  • Comfort operating in a 24/7 production environment where storage incidents directly impact GPU availability, with a security-first approach baked into every build.
  • 100% hands-on engineering. You write and review production code, read file system, kernel, NVMe-oF, or SPDK source to chase bugs, and run scale tests or recovery drills yourself rather than delegating.

Ways to stand out from the crowd:

  • Maintainers or sustained contributions to widely used public projects.
  • Experience crafting or operating storage for AI training or inference at very large GPU scale, with measurable gains in GPU utilization or reductions in I/O bottlenecks.
  • Kernel and file system development experience, metadata scalability, data placement, failure recovery, or HSM or equivalent experience.
  • Kubernetes and CSI driver development for storage.
  • Hands-on experience with SPDK, libfabric, or FUSE performance optimization.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 5, and 272,000 USD - 431,250 USD for Level 6.

About the company

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. Our invention serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is seeking exceptional individuals like you to help us drive the next wave of artificial intelligence.

NVIDIA is widely considered one of the world’s most desirable employers in technology. We have some of the world’s most forward-thinking and passionate people working for us. If you’re creative and autonomous, we want to hear from you!

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

1:47 min

Comparing Egeria to alternative open metadata solutions

Ferd Scheepers · World Congress 2022

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann +3 · LIVE

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

2:08 min

Creating standard APIs via the Egeria open metadata project

Ferd Scheepers · World Congress 2022

3:44 min

Automating storage savings with S3 intelligent tiering

Sébastien Stormacq · World Congress 2021

Videos

See all

Related articles

See all