Principal Storage Architect & Team Lead (Research & HPC Data Platforms)

Stanford University
Stanford, CA, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Compensation
$225,000.0 - $245,000.0
Working hours
Shift work
Job source

Tech stack

Artificial Intelligence Systems Engineering Extract Transform Load (ETL) Ethernet InfiniBand Linux Kernel High Performance Computing Storage Technologies Data Management ZFS File System

Job description

  • Technical Leadership: Lead and mentor a specialized team of systems engineers, balancing high-level architectural design with hands-on operations and escalation support.
  • Tiered Storage Architecture: Oversee the integration of Lustre HSM on the Elm platform, managing data movement policies between parallel filesystems and MinIO object storage.
  • Platform Ownership: Drive the scaling, reliability, security, compliance, operations, and lifecycle management of our primary research computing storage platforms, including for high-risk data.
  • Performance Engineering: Tune I/O for large-scale High Performance Computing and AI workloads.
  • Community Stewardship: Represent Stanford within the Lustre community and other key community groups, contributing to the upstream roadmap and maintaining a vendor-neutral storage strategy., * Interpersonal Skills: Demonstrates the ability to work well with Stanford colleagues and clients and with external organizations.
  • Promote Culture of Safety: Demonstrates commitment to personal responsibility and value for safety; communicates safety concerns; uses and promotes safe behaviors based on training and lessons learned.
  • Subject to and expected to stay in sync with all applicable University policies and procedures, including but not limited to the personnel policies and other policies found in Stanford’s Administrative Guide, http://adminguide.stanford.edu ., The job duties listed are typical examples of work performed by positions in this job classification and are not designed to contain or be interpreted as a comprehensive inventory of all duties, tasks, and responsibilities. Specific duties and responsibilities may vary depending on department or program needs without changing the general nature and scope of the job or level of responsibility. Employees may also perform other duties as assigned.

Requirements

  • Education: Bachelor’s degree and ten years on increasingly technical work experience or a combination of education and relevant experience.
  • Expertise at Scale: 10+ years of hands-on experience architecting, building, and managing Lustre and ZFS or similar filesystems at the 20PB+ scale.
  • Management: Proven experience leading technical teams in a High-Performance Computing (HPC) or Research Computing environment.
  • Object Storage & HSM: Deep technical fluency in MinIO and Lustre HSM (copytools, policy engines like RobinHood) or similar tools.
  • Kernel & Network Mastery: Expert-level knowledge of the Linux kernel and large-scale InfiniBand/Ethernet fabric tuning.
  • ā€œHands-Onā€ Requirement: Must be comfortable in the ā€œweedsā€-capable of debugging issues such as kernel panics, LNet congestion, and metadata bottlenecks alongside the team.

Physical Requirements*:

  • Constantly perform desk-based computer tasks.
  • Frequently sit, grasp lightly/fine manipulation.
  • Occasionally stand/walk, writing by hand.
  • Rarely use a telephone, lift/carry/push/pull objects that weigh up to 10 pounds.

  • Consistent with its obligations under the law, the University will provide reasonable accommodation to any employee with a disability who requires accommodation to perform the essential functions of the job.

Benefits & conditions

  • May work extended hours, evenings, and weekends., The expected pay range for this position is $225,000 to $245,000 per annum.

Stanford University provides pay ranges representing its good faith estimate of the salary or hourly wage the university reasonably expects to pay for a position upon hire. The pay offered to a selected candidate will be determined based on factors such as (but not limited to) the scope and responsibilities of the position, the qualifications of the selected candidate, departmental budget availability, internal equity, geographic location, and external market pay for comparable jobs. At Stanford University, base pay represents only one aspect of the comprehensive rewards package.

At Stanford University, base pay represents only one aspect of the comprehensive rewards package. The Cardinal at Work website ( https://cardinalatwork.stanford.edu/benefits-rewards ) provides detailed information on Stanford’s extensive range of benefits and rewards offered to employees. Specifics about the rewards package for this position may be discussed during the hiring process.

About the company

Stanford University operates one of the most sophisticated academic storage ecosystems in the world, with aggregate capacity exceeding 100PB and 5 billion files. We are seeking a world-class technical leader to oversee our primary research storage platforms. These services span the gamut from the 15PB flash-based Lustre scratch filesystem on our Sherlock HPC cluster to archival storage on the Elm HSM platform.

Why Stanford? You aren’t just managing a storage cluster; you are architecting a data and storage ecosystem that supports Nobel-caliber research across all disciplines.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:57 min

Routing cross-rack traffic seamlessly with NCCL

Kevin Klues Kevin Klues Ā· World Congress 2025

4:52 min

Connecting namespaces with local virtual ethernet pairs

Oliver Seitz Oliver Seitz Ā· World Congress 2025

1:15 min

Overcoming the challenges of modifying Linux kernel code

Ayesha Kaleem Ā· World Congress 2023

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 Ā· Coffee With Developers

1:51 min

Bypassing the CPU stack with remote direct memory access

Lerna Ekmekcioglu Lerna Ekmekcioglu Ā· Europe 2026 Virtual

2:12 min

Implementing automotive ethernet and connected remote vehicle telemetry applications

David Romić Ā· World Congress 2023

Videos

See all

Related articles

See all