Slingshot Hardware System Administrator
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+1 more
Job description
Experteer Overview In this hybrid role, you will manage hardware and software for Slingshot-based supercomputers, ensuring high uptime and robust validation cycles. You’ll work with ASIC, hardware, and software teams to support silicon bring-up and development workloads. You will maintain software stacks, automate imaging, and improve system monitoring while guiding colleagues. This position offers hands-on HPC exposure and impact on cutting-edge networking ASIC validation at HPE. Compensation / Benefits * Debug prototype and Slingshot system issues to maximize uptime * Collaborate with ASIC, hardware, and software teams for silicon validation and bring-up testing * Update development software stack on systems daily * Maintain and improve system status reporting infrastructure * Debug to isolate root causes of failures * Build compute node images with current OS versions and software * Administer compute cluster for the validation team * Communicate progress and issues with management and internal partners * Provide guidance to less-experienced staff Tasks * Bachelor’s or Master’s degree in Computer Science, Information Systems, or equivalent * Typically 4-6 years as a system and/or network administrator and lab technician * Experience with Linux servers, filesystems, and networks * Experience with batch compute clusters (LSF, Slurm) * Scripting ability in Shell/Bash and Python * Hardware lab environment experience and safe electrical procedures * Strong collaboration and communication skills; fast learner Key requirements * Health & wellbeing benefits * Personal & professional development programs * Unconditional inclusion and inclusive culture
Requirements
- and internal partners * Provide guidance to less-experienced staff Tasks * Bachelor’s or Master’s degree in Computer Science, Information Systems, or equivalent * Typically 4-6 years as a system and/or network administrator and lab technician * Experience with Linux servers, filesystems, and networks * Experience with batch compute clusters (LSF, Slurm) * Scripting ability in Shell/Bash and Python * Hardware lab environment experience and safe electrical procedures * Strong collaboration and communication skills; fast learner Key requirements * Health & wellbeing benefits * Personal & professional development programs * Unconditional inclusion and inclusive culture
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Top 6 Hackathons for Developers in 2023
9 Ways to Make Money Hacking
Why Upskilling And Reskilling is Important For Developers
Long-Term Employment vs Job Hopping