> Markdown version of [/jobs/ext/2248880-sr-hpc-system-administrator](https://www.wearedevelopers.com/jobs/ext/2248880-sr-hpc-system-administrator). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. HPC System Administrator - **Company:** The University of Chicago - **Location:** Chicago, IL, United States - **Experience:** Expert - **Salary:** $100,300.0 - $125,000.0 - **Contract:** Permanent contract - **Skills:** Microsoft Windows, Software Applications, Computer Clusters, Compilers, Program Optimization, Computer Programming, Computer Networks, System Configuration, Linux, Distributed File Systems, File Systems, Distributed Computing Environment, General Parallel File Systems, InfiniBand, Intrusion Detection and Prevention, Subnetting, Linux System Administration, Linux Servers, OpenMP, Performance Tuning, Queue Management Systems, Software Tools, Ansible, Server Administration, Utility Software, Network Switches, Network Storage, Firewalls (Computer Science), Information Technology, Deployment Automation, Performance Monitor, Patch Management, Free and Open-Source Software, Slurm, Puppet - **Published:** August 26, 2026 - **Apply:** https://uchicago.wd5.myworkdayjobs.com/External/job/Chicago-IL/Sr-HPC-System-Administrator_JR34672 ## About the Role Minimum requirements include a college or university degree in related field. Work Experience: Minimum requirements include knowledge and skills developed through 5-7 years of work experience in a related job discipline., * Master's degree in Computer Science or closely related field. Experience: * Full time Linux system administration experience in a large distributed computing environment. * Previous experience in providing support for Linux HPC cluster used for scientific research. Technical Skills or Knowledge: * Experience with installing, configuring, and maintaining job management tools (such as SLURM, Moab, TORQUE, PBS, etc.). * Experience configuring, installing and troubleshooting MPI and OpenMP. * Experience with operating system deployment tools (e.g. XCAT, ROCKS). * Experience configuring, administering, and supporting network storage subsystems (e.g. IBM, NetAppl DataDirect Network, LSI, etc.). * Hands-on experience of at least one distributed file system (Spectrum Scale-GPFS, Lustre, BeeGFS, Gluster, IMRIX, PVFS, etc.). * Direct experience working with Infiniband (must at least be able to demonstrate a working knowledge of Infiniband concepts, OFED layers, sub-net managers). * Experience configuring, installing, tuning and maintaining scientific application software on large-scale systems. * Experience supporting HPC compilers and libraries. * Experience with systems automation tools such as Ansible or Puppet. * Experience configuring, installing, maintaining and/or using performance monitoring and optimization tools. Preferred Competencies * Ability to work well with faculty and researchers. * Ability to identify and gain expertise in appropriate new technologies and/or software tools. * Ability to function as part of an interactive team while demonstrating self-initiative to achieve project's goals and Research Computing Center's mission. * Strong analytical skills and problem-solving ability. ## Description The job uses specialized knowledge and breadth of expertise to design automated, scalable, and rapidly deployable solutions to infrastructure development and server configuration. Leads installation, configuration, and maintenance of operating systems. Uses best practices and systems knowledge to monitor and alert systems, utility software, and firewalls. Guides maintenance for production servers as well as Windows and Linux servers., The University of Chicago is seeking a highly qualified Senior HPC System Administrator to join the system and operation team that builds and manages RCC HPC systems and facility operations. The individual in this position will be involved in the procurement and management of HPC hardware and software., * Installing, configuring, and maintaining large computer clusters/servers and software. * Day-to-day operations of the systems including systems administration, monitoring and storage performance up to and including network components. Management of the system's network switch, parallel file system and HPC software stack and tools. * Configuration of the scheduling and queuing system. * Diagnosing and resolving system operational problems quickly and effectively. Coordinating with vendors to resolve hardware and software problems. Assist users with access and other help desk ticket requests or issues. * Use scripting/programming skills to enable system-level automation, problem detection, security maintenance and patch management. * Building and deploying open-source software and software from vendors/partners. * Providing reliable and efficient backups/restores for all managed systems. * Documenting system administration procedures for routine and complex tasks. * Maintaining and monitoring the security of the HPC systems and servers. * Plans and installs necessary patches and upgrades for servers and their associated storage, network, communications, and peripheral sub-systems. Installs and maintains an appropriate level of intrusion detection, monitoring, and auditing software as required. * Tracks compliance and maintains documentation for hardware, software, and service inventories for management reports. * Performs other related work as needed. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Automate everything via NodeJS and Puppeteer](https://www.wearedevelopers.com/videos/322-automate-everything-via-nodejs-and-puppeteer) - [Schroedinger's cat: Thinking in- and outside the box of quantum mechanics](https://www.wearedevelopers.com/videos/216-schroedinger-s-cat-thinking-in-and-outside-the-box-of-quantum-mechanics) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [The Quantum Computing Future](https://www.wearedevelopers.com/videos/1145-the-quantum-computing-future) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [Best Coding Boot Camps in Germany](https://www.wearedevelopers.com/magazine/237-best-coding-boot-camps-in-germany) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [7 Important Tips That Every Software Developer Should Know](https://www.wearedevelopers.com/magazine/101-7-important-tips-that-every-software-developer-should-know) - [Should senior developers refuse interview coding challenges?](https://www.wearedevelopers.com/magazine/29-should-senior-developers-refuse-interview-coding-challenges)