HPC Systems Administrator (Hardware & Infrastructure Operations)

Stanford University
Stanford, CA, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$150,289.0 - $171,674.0
Working hours
Shift work
Job source

Tech stack

Data Centers Data Center Infrastructure Management (CIM) Ethernet Firmware Job Scheduling Linux System Administration Log Analysis Scripting Infrastructure Automation Frameworks Slurm Hardware Asset Management Hardware Infrastructure

Job description

  • Hardware Lifecycle & Deployment: Leadthephysicaldeployment,burn-in, troubleshooting,anddecommissioningofcomputenodes,GPUservers,and high-density storage systems.
  • Diagnostics & Root Cause Analysis: Perform troubleshooting on hardware issues-suchasmemoryerrors,GPUthermalthrottling,networkfailures-and coordinate with vendors for support and replacements.
  • Data Center Operations: Collaboratewiththedatacentersteamtoplanandmanage hardware deployments.
  • Provisioning & Automation: Work with lead platform administrators on testing and provisioningtoensurerapid,consistentdeploymentofclusterimagesacrossthefleet.
  • Health & Telemetry: Refinehardware-levelmonitoringtoproactivelyidentifyfailing components before they impact active research jobs.

Requirements

  • Education: Bachelor’s degree and eight years of relevant experience, or a combination of education and relevant experience.
  • Experience: 3-5+ years of experience in Linux Systems Administration, with a strong preference for candidates from HPC, larges-scale data center, or research environments.
  • Hardware Proficiency: Solid understanding of x86 server architecture, GPU systems, ethernet,and high-performance interconnects.
  • Scripting: Proficiency in scripting languages for automating hardware health checks, log parsing, and routine maintenance tasks.
  • Infrastructure Management: Experience using configuration management tools to manage hardware settings and firmware versions at scale. Experience working with data center teams to populate and maintain DCIM solutions preferred.
  • Physical Requirements: Ability to lift up to 50 lbs and work comfortably in a data center environment, including racking equipment and managing complex cable topologies.
  • Communication: Strong written and verbal communication skills.

Preferred Skills

  • Direct experience maintaining hardware for HPC systems and large scale storage systems.

  • Familiarity with the Slurm workload manager and how hardware health impacts job scheduling.

  • Exposure to liquid cooling solutions or high-density rack power management.

Physical Requirements*:

  • Constantly perform desk-based computer tasks.

  • Frequently sit, grasp lightly/fine manipulation.

  • Occasionally stand/walk, writing by hand.

  • Rarely use a telephone, lift/carry/push/pull objects that weigh up to 10 pounds., * Interpersonal Skills: Demonstrates the ability to work well with Stanford colleagues and clients and with external organizations.

  • Promote Culture of Safety: Demonstrates commitment to personal responsibility and value for safety; communicates safety concerns; uses and promotes safe behaviors based on training and lessons learned.

Benefits & conditions

  • May work extended hours, evenings, and weekends.

About the company

The Sherlock HPC cluster is the flagship of Stanford’s research computing environment, supporting thousands of users and a massive variety of scientific workloads. We are looking for an HPC Systems Administrator who thrives at the intersection of high-density hardware and Linux systems engineering.

In this role, you will be the primary steward of the physical infrastructure on Sherlock and other platforms. You will ensure that our 1,500+ compute nodes, high-density GPU racks, and petabyte-scale storage arrays are meticulously maintained, expertly tuned, and highly available.

Why Stanford?

You won’t just be swapping parts; you will be managing the physical backbone of a world-class research environment. From debugging errors on NVIDIA H200s to optimizing InfiniBand cabling for our Lustre scratch tiers, your work is the foundation upon which Nobel-caliber research is built.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · WWC Europe 2026

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

4:52 min

Connecting namespaces with local virtual ethernet pairs

Oliver Seitz Oliver Seitz · WWC 2025

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac · WWC 2021

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · WWC Europe 2026

2:12 min

Implementing automotive ethernet and connected remote vehicle telemetry applications

David Romić · WWC 2023

Videos

See all

Related articles

See all