> Markdown version of [/jobs/ext/3045235-linux-hpc-storage-engineer](https://www.wearedevelopers.com/jobs/ext/3045235-linux-hpc-storage-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Linux HPC Storage Engineer - **Company:** ORNL FCU - **Location:** Oak Ridge, TN, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Big Data, Ubuntu (Operating System), CentOS, Computer Clusters, Configuration Management, Information Systems, Computer Engineering, Linux, RAID, File Systems, Perl (Programming Language), Ganglia (Software), General Parallel File Systems, Monitoring of Systems, InfiniBand, Storage Area Network (SAN), Job Scheduling, Python (Programming Language), Linux System Administration, Nagios, NetApp Applications, Red Hat Enterprise Linux, Ansible, Tape Libraries, Weka, Zabbix, Ceph (Software), Scripting, Data Storage Management, High Performance Computing, CheckMK, Grafana, Git, Containerization, Infrastructure Automation Frameworks, Storage Technologies, Information Technology, Performance Monitor, Data Management, Slurm, ZFS File System, Puppet, Jenkins, Nvme, Vmware - **Published:** September 24, 2026 - **Apply:** https://www.thejobnetwork.com/job/9265dce9-0ef4-43e1-b410-41b95f0515e9/senior-linux-hpc-storage-engineer ## About the Role * A BS degree in computer science, computer engineering, information technology, information systems, science, engineering, or related discipline and 8-12 years of relevant professional experience; or an equivalent combination of education and experience. + Master's degree holders: 7-10 years of relevant experience. + PhD holders: 4-6 years of relevant experience. * Five (5) or more years managing UNIX/Linux systems. * Demonstrated experience managing HPC storage and large-scale enterprise storage systems. * Three (3) or more years working with configuration management and automation tools such as Git, Jenkins, Ansible, or Puppet. * Proficiency with at least one scripting language (Bash, Python, Perl, etc.). * Strong Linux administration and advanced troubleshooting experience. * Experience supporting large data systems and/or HPC scientific workloads. * Strong desire to innovate and evaluate new technologies for HPC and storage environments. * Collaborative approach and ability to become a trusted advisor to research teams. Preferred Qualifications * Active DOE Q, DoD Top Secret, or TS/SCI clearance is strongly preferred. * Solid understanding of multiple operating systems and HPC cluster technologies. * Experience with Rocky/CentOS/RHEL, Ubuntu, VMware. * Understanding of HPC job schedulers (SLURM) and user support workflows. * Experience with container technologies in HPC environments. * Experience with multiple system deployment mechanisms (Warewulf, PXEboot, Cobbler, Bright). * Experience with GPU clusters (NVIDIA, AMD) for AI/ML and scientific workloads. * Deep expertise with high-performance parallel file systems (Lustre, GPFS/Spectrum Scale, BeeGFS, WEKA). * Knowledge of storage networking (Infiniband, NVMe-oF, SAN/NAS architectures). * Familiarity with RAID, ZFS, and object storage technologies. * Strong background in performance monitoring, benchmarking, and I/O optimization. * Experience with monitoring systems such as Grafana, CheckMK, Nagios, Zabbix, Ganglia. * Previous experience working in a government, scientific, or other highly technical environment. * Strong documentation skills and ability to prepare web-based documentation. ## Description * Must be able to work a hybrid work schedule in Oak Ridge, TN * Must be eligible for a federal security clearance (US Citizen) Major Duties/Responsibilities 1. Architect, deploy, and manage large-scale HPC storage systems, including parallel file systems such as Lustre, GPFS/Spectrum Scale, BeeGFS and WEKA 2. Design, implement, and operate large-scale Ceph storage clusters for HPC and research workloads, delivering reliable, high-performance object, block, and file storage services. 3. Ensure the availability, performance, scalability, and security of production storage environments. 4. Administer and optimize enterprise storage platforms such as Qumulo and NetApp in support of HPC and research workloads. 5. Design, deploy, and maintain archival storage solutions including Spectra Logic BlackPearl and large-scale tape libraries to ensure long-term data preservation and accessibility. 6. Integrate high-performance, enterprise, and archival storage layers into cohesive tiered storage architectures that balance cost, scalability, and performance for diverse scientific workflows. 7. Leverage automation and monitoring solutions to minimize day-to-day maintenance while identifying opportunities to optimize system performance and management. 8. Collaborate with researchers and technical POCs to support large data workflows and optimize I/O performance for scientific workloads. 9. Automate storage provisioning, monitoring, and maintenance using scripting and configuration management tools. 10. Diagnose and resolve complex storage and I/O-related issues in high-throughput, low-latency HPC environments. 11. Evaluate emerging storage technologies (NVMe, object storage, hierarchical storage management, burst buffers) and contribute to strategic planning for future HPC systems. 12. Work with 24/7 operations staff to streamline monitoring and troubleshooting, significantly reducing the need for off-hours support. 13. Deliver ORNL's mission by aligning behaviors, priorities, and interactions with our core values of Impact, Integrity, Teamwork, Safety, and Service. Promote equal opportunity by fostering a respectful workplace. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [How I saved 200K/yr in direct costs writing 0 code lines in K8s](https://www.wearedevelopers.com/videos/1055-how-i-saved-200k-yr-in-direct-costs-writing-0-code-lines-in-k8s) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Stop Committing Your Secrets - GIt Hooks To The Rescue!](https://www.wearedevelopers.com/videos/573-stop-committing-your-secrets-git-hooks-to-the-rescue) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know)