HPC Engineer-2
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+6 more
Job description
-
System Design and Implementation: HPC Engineers Design, Build and configure HPC clusters including hardware and software components
-
System administration: Manage and maintain the HPC infrastructure including operating systems, storage, and networking
-
Performance Optimization: analyze system performance, identify bottlenecks, and implement solutions to optimize performance for various applications
-
Troubleshoot and support: Diagnose and resolve issues with the HPC system, providing support to researchers and users
-
Scripting and Automating: Develop scripts and automation tools to streamline routine tasks and improve efficiency
-
Collaboration: Work closely with researchers, data scientists, and other engineers to understand their needs and provide effective solutions
-
Documentation - Maintain clear and accurate documentation of system configurations, procedures and troubleshooting steps
-
Monitor and Maintenance - Monitor system health, perform maintenance tasks, and plan for upgrades and new technologies
-
Security: Ensure the secure and effective operation of HPC
Requirements
-
Experience in building and deploying scientific applications and module environments in on-premises and cloud-based HPC environments.
-
Familiarity with Open OnDemand, AWS cloud-based infrastructure, and containerization of HPC applications (e.g., Singularity Registry HPC)
-
Knowledge of the PBS Grid Engine, Slurm job scheduler.
-
Excellent troubleshooting skills with the ability to resolve application-related issues.
-
Strong documentation and diagramming abilities.
-
Ability to work collaboratively within a team and communicate effectively.
-
Linux Administration
-
Linux OS internal knowledge
-
Linux Hardening
-
LDAP / AD intergrations
-
Bash Scripting
Preferred Qualifications:
-
Ansible Configuration Management
-
Splunk and Nagios monitoring
-
Base Command Manager (formerly called Bright Cluster Manager)
-
Jira
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
7 Cloud Computing Trends Coming in 2025 for Developers
Top 6 Hackathons for Developers in 2023
Dev Digest 121 - AI goes offline
Highest Paying Tech Companies for Developers