HPC Administrator

Atos SE
Madrid, Spain
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Bash Shell Computer Clusters Nvidia CUDA Linux File Systems Ethernet Firmware Monitoring of Systems InfiniBand Python (Programming Language) Linux System Administration OpenMP
+11 more
Open Source Technology Performance Tuning Software Maintenance Red Hat Enterprise Linux Scientific Computating Scripting High Performance Computing Parallel Computation Build Management Low Latency Slurm

Job description

You will support the design, build, and operation of Linux-based high-performance computing clusters (CPU and GPU). You will manage system administration tasks, ensure cluster stability, maintain software stacks, and provide support for HPC users and scientific computing environments., * Design and build Linux-based HPC CPU/GPU clusters.

  • Perform system administration, software maintenance, system monitoring, and troubleshooting for HPC/GPU clusters.
  • Manage and operate HPC facilities and infrastructure.
  • Administer Red Hat Enterprise Linux 8, 9, 10 or similar operating systems.
  • Configure and support parallel computing environments CUDA, OpenMP, MPI.
  • Deploy and manage Infiniband (ultra low latency networks) and Ethernet high-performance networks.
  • Administer batch scheduling systems such as Slurm.
  • Manage compute resources (CPU, GPU, RAM) for serial and parallel workloads.
  • Compile, install, update, and tune scientific software and libraries (commercial and open-source) on compute nodes and HPC file systems.
  • Maintain and update firmware and drivers related to HPC systems.
  • Provide user support for HPC environments, including troubleshooting, software assistance, guidance on cluster usage, and performance optimization.

Requirements

  • Experience designing and managing HPC Linux clusters (CPU/GPU).
  • Strong background in Linux system administration (preferably RHEL-based).
  • Knowledge of parallel programming environments (CUDA, OpenMP, MPI).
  • Experience administering Slurm and managing HPC compute resources.
  • Hands-on experience with Infiniband and high-performance networking.
  • Proficiency with compiling and maintaining scientific software stacks.
  • Strong troubleshooting skills in complex HPC environments.

Nice to Have

  • Familiarity with HPC storage systems.
  • Scripting knowledge (Bash, Python).
  • Experience in performance tuning of HPC environments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jobleads.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:37 min

Simplifying parallel programming with the CUDA ecosystem

Paul Graham Paul Graham · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · WWC Europe 2026

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

1:11 min

Running high-performance edge computing on bare metal servers

Josip Stuhli Josip Stuhli · Coffee With Developers

Videos

See all

Related articles

See all