> Markdown version of [/jobs/ext/1747605-hpc-data-storage-administrator-research-hpc-data-platforms](https://www.wearedevelopers.com/jobs/ext/1747605-hpc-data-storage-administrator-research-hpc-data-platforms). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # HPC Data Storage Administrator (Research & HPC Data Platforms) - **Company:** Stanford University - **Location:** Stanford, CA, United States - **Experience:** Expert - **Salary:** $150,289.0 - $171,674.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computing Platforms, Big Data, Extract Transform Load (ETL), Data Loss, Data Security, Software Debugging, Ethernet, Firmware, InfiniBand, Linux Kernel, Linux System Administration, Parsing, Performance Tuning, Scripting, Data Storage Technologies, High Performance Computing, Storage Technologies, Data Management, ZFS File System - **Published:** July 6, 2026 - **Apply:** https://dejobs.org/x/x/F8DD1DE032C9421A93E720591EA6AC5F/job/ ## About the Role * Education: Bachelor's degree and eight years of relevant experience, or a combination of education and relevant experience. * Expertise at Scale: 8+ years of hands-on experience architecting, building, and managing Lustre and ZFS or similar filesystems at the 20PB+ scale. * Object Storage & HSM: Deep technical fluency in MinIO and Lustre HSM (copytools, policy engines like RobinHood) or similar tools. * Kernel & Network Mastery: Expert-level knowledge of the Linux kernel and large-scale InfiniBand/Ethernet fabric tuning. * In-depth Troubleshooting Experience: Must be capable of leading the debugging of issues such as kernel panics, LNet congestion, and metadata bottlenecks. * Leadership: Proven experience mentoring junior admins and leading large-scale migration projects without data loss. * Communication: Strong written and verbal communication skills., * Experience: 5+ years of Linux Systems Administration, with 3+ years specifically in an HPC or large-scale data environment. * Technical Stack: Strong hands-on experience with Lustre, ZFS, MinIO, and/or similar technologies. * Scripting: Advanced proficiency in scripting languages for automating routine storage tasks and parsing system logs. * Hardware Mastery: Comfortable with the physical aspects of the role-diagnosing hardware failures and understanding power/cooling requirements for high-density storage. * Communication: Strong written and verbal communication skills. Qualifications: Physical Requirements*: * Constantly perform desk-based computer tasks. * Frequently sit, grasp lightly/fine manipulation. * Occasionally stand/walk, writing by hand. * Rarely use a telephone, lift/carry/push/pull objects that weigh up to 10 pounds. * Consistent with its obligations under the law, the University will provide reasonable accommodations to applicants and employees with disabilities. Applicants requiring a reasonable accommodation for any part of the application or hiring process should contact Stanford University Human Resources by submitting a contact form (https://stanford.service-now.com/humanresources_services?id=sc_cat_item&sys_id=aa4161da130d574019813598d144b0b2) . ## Description * Architecture: Adapt and evolve the technology designs of existing systems to meet the needs of future computing platforms and research aims. * Platform Management: Deliver on the scaling, reliability, security, compliance, operations, and lifecycle management of our primary research computing storage platforms, including for high-risk data. * Tiered Storage Architecture: Oversee the integration of Lustre HSM on the Elm platform, managing data movement policies between parallel filesystems and MinIO object storage. * Performance Engineering: Tune I/O for large-scale High Performance Computing and AI workloads. * Community Stewardship: Represent Stanford within the Lustre community and other key community groups, contributing to the upstream roadmap and maintaining a vendor-neutral storage strategy., * Platform Management: Contribute to the scaling, reliability, security, compliance, operations, and lifecycle management of our primary research computing storage platforms, including for high-risk data. * Operational Excellence: Perform complex filesystem upgrades, kernel patches, and hardware refreshes with minimal downtime. * Monitoring & Telemetry: In collaboration with others, build and maintain sophisticated observability stacks for real-time I/O tracking and trend analysis. * User Support: Act as an escalation point for researchers struggling with complex I/O patterns, job failures, or data access issues. * Maintenance: Manage the physical and logical health of the storage fleet, including RMA processes, firmware updates, and disk replacement cycles., * Interpersonal Skills: Demonstrates the ability to work well with Stanford colleagues and clients and with external organizations. * Promote Culture of Safety: Demonstrates commitment to personal responsibility and value for safety; communicates safety concerns; uses and promotes safe behaviors based on training and lessons learned. * Subject to and expected to stay in sync with all applicable University policies and procedures, including but not limited to the personnel policies and other policies found in Stanford's Administrative Guide, http://adminguide.stanford.edu., The job duties listed are typical examples of work performed by positions in this job classification and are not designed to contain or be interpreted as a comprehensive inventory of all duties, tasks, and responsibilities. Specific duties and responsibilities may vary depending on department or program needs without changing the general nature and scope of the job or level of responsibility. Employees may also perform other duties as assigned. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Tips and Tricks for Working with JSON](https://www.wearedevelopers.com/videos/1229-tips-and-tricks-for-working-with-json) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) ## Related Articles - [What does the history of data storage tell us about the future?](https://www.wearedevelopers.com/magazine/495-what-does-the-history-of-data-storage-tell-us-about-the-future) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story)