HPC System Administrator - Hybrid
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+25 more
Job description
An HPC Systems Administrator is responsible for managing and maintaining high-performance computing (HPC) systems, including hardware, software, and network infrastructure, ensuring optimal performance and stability for computationally intensive research or business applications, often involving tasks like installing and configuring specialized software, monitoring system health, optimizing job scheduling, and troubleshooting complex technical issues related to large-scale computing environments., * System Installation and Configuration:
- Installing and configuring HPC hardware (clusters, servers, storage)
- Deploying and configuring HPC software (operating systems, job schedulers, parallel programming environments)
- Setting up network infrastructure for high-speed data transfer
- Performance Optimization:
- Monitoring system health and resource utilization
- Analyzing job performance and identifying bottlenecks
- Optimizing job scheduling algorithms to maximize throughput
- Tuning system parameters for optimal performance
- User Support:
- Providing technical support to HPC users on system access, job submission, and troubleshooting
- Developing user documentation and training materials
- Security Management:
- Implementing security policies and procedures for HPC systems
- Managing user access and permissions
- Monitoring for potential security threats
- Backup and Disaster Recovery:
- Implementing data backup strategies for HPC systems
- Testing and maintaining disaster recovery plans
- Hardware Maintenance:
- Coordinating hardware repairs and upgrades with vendors, * All job-specific, safety, and compliance training will be assigned based on the job functions associated with this employee.
Other
- This position requires periodic travel and some evenings, weekends, and/or holidays.
- Job may require after-hours response to emergency issues.
- Periodically scheduled on-call may require after-hours response for technical emergencies not explicitly related to assigned job responsibilities
Requirements
- Bachelorâs degree in Computer Science, Information Technology, Engineering, or a related technical field, or three (3) years of directly related Linux systems administration or HPC infrastructure experience in lieu of a bachelorâs degree.
- Minimum of 3 years of Linux system administration experience in an enterprise or HPC environment.
- Minimum of 2 years administering or supporting High Performance Computing (HPC) infrastructure.
- Minimum of 2 years of experience with HPC job schedulers such as IBM LSF
- Minimum of 2 years of experience administering parallel file systems, such as IBM Spectrum Scale (GPFS), Lustre, or BeeGFS.
- Minimum of 2 years supporting enterprise storage solutions, such as NetApp E-Series or similar SAN storage platforms.
- Experience supporting high speed networking technologies, including InfiniBand or equivalent low latency networking.
- Experience troubleshooting Linux operating systems, networking, storage, and distributed computing environments.
- Experience with shell scripting (Bash) and/or Python for system administration and automation.
- Experience installing, configuring, and maintaining enterprise server hardware in a data center environment.
- Ability to participate in an on-call rotation and respond to after-hours production incidents.
- Ability to lift and install server, storage, and networking equipment as required.
- Proficient in Microsoft Office Suite, specifically Word, Excel, Outlook, and general working knowledge of Internet for business use, * Basic understanding of Linux operating system.
- Expertise in HPC software and hardware systems like IBM Spectra Scale, NetApp E-Series storage, Nvidia InfiniBand and job schedulers (e.g., LSF, Slurm)
- Knowledge of network administration and high-performance networking protocols
- Experience with system monitoring and performance analysis tools
- Excellent problem-solving and troubleshooting abilities
Physical Demands
- Ability to lift, move and install HPC data center hardware and supplies.
- Standing for extended periods while performing data center related tasks.
About the company
At Caris, we understand that cancer is an ugly word-a word no one wants to hear, but one that connects us all. Thatâs why weâre not just transforming cancer care-weâre changing lives.
We introduced precision medicine to the world and built an industry around the idea that every patient deserves answers as unique as their DNA. Backed by cutting-edge molecular science and AI, we ask ourselves every day: âWhat would I do if this patient were my mom?â That question drives everything we do.
But our mission doesnât stop with cancer. Weâre pushing the frontiers of medicine and leading a revolution in healthcare-driven by innovation, compassion, and purpose.
Join us in our mission to improve the human condition across multiple diseases. If youâre passionate about meaningful work and want to be part of something bigger than yourself, Caris is where your impact begins., Caris Life Sciences is a leading innovator in molecular science and artificial intelligence focused on fulfilling the promise of precision medicine through quality and innovation.
Caris is committed to quality and excellence at our state-of-the-art laboratories. Learn more about our tissue lab and the advanced technologies that are helping improve the lives of cancer patients.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on dejobs.orgGood distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
7 Cloud Computing Trends Coming in 2025 for Developers
A Guide to Green Tech and Green IT Careers
Best Paying Jobs in Technology