HPC Administrator
Atos SE
Madrid, Spain
about 1 month ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source
Tech stack
Bash Shell
Computer Clusters
Nvidia CUDA
Linux
File Systems
Ethernet
Firmware
Monitoring of Systems
InfiniBand
Python (Programming Language)
Linux System Administration
OpenMP
+11 more
Open Source Technology
Performance Tuning
Software Maintenance
Red Hat Enterprise Linux
Scientific Computating
Scripting
High Performance Computing
Parallel Computation
Build Management
Low Latency
Slurm
Job description
You will support the design, build, and operation of Linux-based high-performance computing clusters (CPU and GPU). You will manage system administration tasks, ensure cluster stability, maintain software stacks, and provide support for HPC users and scientific computing environments., * Design and build Linux-based HPC CPU/GPU clusters.
- Perform system administration, software maintenance, system monitoring, and troubleshooting for HPC/GPU clusters.
- Manage and operate HPC facilities and infrastructure.
- Administer Red Hat Enterprise Linux 8, 9, 10 or similar operating systems.
- Configure and support parallel computing environments CUDA, OpenMP, MPI.
- Deploy and manage Infiniband (ultra low latency networks) and Ethernet high-performance networks.
- Administer batch scheduling systems such as Slurm.
- Manage compute resources (CPU, GPU, RAM) for serial and parallel workloads.
- Compile, install, update, and tune scientific software and libraries (commercial and open-source) on compute nodes and HPC file systems.
- Maintain and update firmware and drivers related to HPC systems.
- Provide user support for HPC environments, including troubleshooting, software assistance, guidance on cluster usage, and performance optimization.
Requirements
- Experience designing and managing HPC Linux clusters (CPU/GPU).
- Strong background in Linux system administration (preferably RHEL-based).
- Knowledge of parallel programming environments (CUDA, OpenMP, MPI).
- Experience administering Slurm and managing HPC compute resources.
- Hands-on experience with Infiniband and high-performance networking.
- Proficiency with compiling and maintaining scientific software stacks.
- Strong troubleshooting skills in complex HPC environments.
Nice to Have
- Familiarity with HPC storage systems.
- Scripting knowledge (Bash, Python).
- Experience in performance tuning of HPC environments.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.jobleads.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
about 3 years ago
LM
Luis Minvielle
7 Cloud Computing Trends Coming in 2025 for Developers
over 2 years ago
BB
Benedikt Bischof
Making Data Warehouses Fast: A Developer’s Story
about 4 years ago
DD
Dilek Demir
Data Science & more: The Lopez dilemma
almost 6 years ago
EM
Eli McGarvie
Highest Paying Tech Companies for Developers
over 3 years ago
CH
Chris Heilmann
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
almost 2 years ago