> Markdown version of [/jobs/ext/1705542-hpc-systems-administrator-hardware-infrastructure-operations](https://www.wearedevelopers.com/jobs/ext/1705542-hpc-systems-administrator-hardware-infrastructure-operations). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # HPC Systems Administrator (Hardware & Infrastructure Operations) - **Company:** Stanford University - **Location:** Stanford, CA, United States - **Experience:** Experienced - **Salary:** $150,289.0 - $171,674.0 - **Contract:** Permanent contract - **Skills:** Data Centers, Data Center Infrastructure Management (CIM), Ethernet, Firmware, Job Scheduling, Linux System Administration, Log Analysis, Scripting, Infrastructure Automation Frameworks, Slurm, Hardware Asset Management, Hardware Infrastructure - **Published:** July 6, 2026 - **Apply:** https://dejobs.org/x/x/F08767A9EBF7413DBA1A9EF29AAB6CF4/job/ ## About the Role * Education: Bachelor's degree and eight years of relevant experience, or a combination of education and relevant experience. * Experience: 3-5+ years of experience in Linux Systems Administration, with a strong preference for candidates from HPC, larges-scale data center, or research environments. * Hardware Proficiency: Solid understanding of x86 server architecture, GPU systems, ethernet,and high-performance interconnects. * Scripting: Proficiency in scripting languages for automating hardware health checks, log parsing, and routine maintenance tasks. * Infrastructure Management: Experience using configuration management tools to manage hardware settings and firmware versions at scale. Experience working with data center teams to populate and maintain DCIM solutions preferred. * Physical Requirements: Ability to lift up to 50 lbs and work comfortably in a data center environment, including racking equipment and managing complex cable topologies. * Communication: Strong written and verbal communication skills. Preferred Skills * Direct experience maintaining hardware for HPC systems and large scale storage systems. * Familiarity with the Slurm workload manager and how hardware health impacts job scheduling. * Exposure to liquid cooling solutions or high-density rack power management. Physical Requirements*: * Constantly perform desk-based computer tasks. * Frequently sit, grasp lightly/fine manipulation. * Occasionally stand/walk, writing by hand. * Rarely use a telephone, lift/carry/push/pull objects that weigh up to 10 pounds., * Interpersonal Skills: Demonstrates the ability to work well with Stanford colleagues and clients and with external organizations. * Promote Culture of Safety: Demonstrates commitment to personal responsibility and value for safety; communicates safety concerns; uses and promotes safe behaviors based on training and lessons learned. ## Description * Hardware Lifecycle & Deployment: Leadthephysicaldeployment,burn-in, troubleshooting,anddecommissioningofcomputenodes,GPUservers,and high-density storage systems. * Diagnostics & Root Cause Analysis: Perform troubleshooting on hardware issues-suchasmemoryerrors,GPUthermalthrottling,networkfailures-and coordinate with vendors for support and replacements. * Data Center Operations: Collaboratewiththedatacentersteamtoplanandmanage hardware deployments. * Provisioning & Automation: Work with lead platform administrators on testing and provisioningtoensurerapid,consistentdeploymentofclusterimagesacrossthefleet. * Health & Telemetry: Refinehardware-levelmonitoringtoproactivelyidentifyfailing components before they impact active research jobs. ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Minimising the Carbon Footprint of Workloads](https://www.wearedevelopers.com/videos/1198-minimising-the-carbon-footprint-of-workloads) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Agent Smith Gets Hardware: Autonomous IoT Hacking From Debug Port to Cloud API](https://www.wearedevelopers.com/videos/100258-agent-smith-gets-hardware-autonomous-iot-hacking-from-debug-port-to-cloud-api) ## Related Articles - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [A Guide to Green Tech and Green IT Careers](https://www.wearedevelopers.com/magazine/374-a-guide-to-green-tech-and-green-it-careers) - [Data Science & more: The Lopez dilemma](https://www.wearedevelopers.com/magazine/10-data-science-more-the-lopez-dilemma) - [Dev Digest 134 - Where pixels sing?](https://www.wearedevelopers.com/magazine/477-dev-digest-134-where-pixels-sing)