Hpc Engineer

Hermès
Zaragoza, Spain
1 day ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

System Configuration Linux Firmware Monitoring of Systems Performance Tuning Scientific Computating User Provisioning Software Scripting High Performance Computing Software Troubleshooting Containerization Infrastructure Automation Frameworks

Job description

We are looking for an experiencedAI & HPC Infrastructure Engineerto operate, optimise, and continuously evolve high-performance computing infrastructure supporting advanced engineering and scientific workloads.This is a hands-on infrastructure role for someone with strong Linux and HPC expertise who is comfortable owning production platforms, troubleshooting complex systems, and driving improvements across GPU compute, storage, networking, containers, and platform performance.What You’ll DoAdminister and maintain GPU compute infrastructure, including system configuration, NVIDIA drivers and software stack, firmware, and hardware health monitoring.Operate, monitor, and optimise HPC clusters to ensure reliability, performance, and efficient use of compute resources.Define and refine resource allocation policies across different workload types to ensure effective and equitable use of available capacity.Manage containerised environments for platform users, including user provisioning, environment maintenance, GPU access, and standards for container usage.Administer shared and high-performance storage and support the operation and troubleshooting of high-speed interconnects.Lead performance engineering activities, including profiling, benchmarking, bottleneck analysis, and platform optimisation.Support the integration and optimisation of engineering and scientific applications on GPU-accelerated and HPC platforms.Provide incident response, troubleshooting, root-cause analysis, and change management, including planning and execution of maintenance activities.Maintain platform monitoring, alerting, and operational reporting.Develop and maintain technical documentation, operational procedures, and platform standards.Contribute to security, architecture governance, and compliance activities.Required QualificationsSubstantial professional experience in Linux systems engineering / administration at an advanced level.Proven experience operating production HPC environments, including cluster administration and large-scale parallel workloads.Hands-on experience administering GPU compute infrastructure, particularly NVIDIA-based environments and the associated software stack.Practical experience with containerisation in HPC environments, including GPU device access and multi-node workloads.Strong troubleshooting and performance-tuning skills across complex compute infrastructure.Demonstrated ability to take end-to-end ownership of production infrastructure and act as a senior technical escalation point.Strong analytical and problem-solving skills with a production-focused mindset.Professional working proficiency in English.Preferred QualificationsExperience supporting engineering simulation, scientific computing, or computational analysis workloadson accelerated infrastructure.Experience with HPC scheduling and resource management.Experience with monitoring and observability for compute infrastructure.Experience with high-performance storage and networking/interconnect technologies.Experience in an enterprise, industrial, or regulated environment.Experience with platform automation, scripting, or infrastructure tooling.Ideal CandidateThe ideal candidate is a hands-on infrastructure engineer who combines deep Linux, HPC, and GPU expertise with strong operational ownership.You are comfortable working close to the hardware and software stack, diagnosing complex performance and reliability issues, and improving infrastructure used by demanding engineering and scientific workloads. You thrive in environments where reliability, performance, and technical depth matter, and you can operate effectively as a senior technical point of reference for the platform.#J-*****-Ljbffr

Requirements

Substantial professional experience in Linux systems engineering / administration at an advanced level. Proven experience operating production HPC environments, including cluster administration and large-scale parallel workloads. Hands-on experience administering GPU compute infrastructure, particularly NVIDIA-based environments and the associated software stack. Practical experience with containerisation in HPC environments, including GPU device access and multi-node workloads. Strong troubleshooting and performance-tuning skills across complex compute infrastructure. Demonstrated ability to take end-to-end ownership of production infrastructure and act as a senior technical escalation point. Strong analytical and problem-solving skills with a production-focused mindset. Professional working proficiency in English. Preferred Qualifications Experience supporting engineering simulation, scientific computing, or computational analysis workloadson accelerated infrastructure. Experience with HPC scheduling and resource management. Experience with monitoring and observability for compute infrastructure. Experience with high-performance storage and networking/interconnect technologies. Experience in an enterprise, industrial, or regulated environment. Experience with platform automation, scripting, or infrastructure tooling. Ideal Candidate The ideal candidate is a hands-on infrastructure engineer who combines deep Linux, HPC, and GPU expertise with strong operational ownership. You are comfortable working close to the hardware and software stack, diagnosing complex performance and reliability issues, and improving infrastructure used by demanding engineering and scientific workloads. You thrive in environments where reliability, performance, and technical depth matter, and you can operate effectively as a senior technical point of reference for the platform. #J-*****-Ljbffr

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac · World Congress 2021

1:04 min

Introduction to Bitcoin script parsing tools

Steve Shadders · LIVE

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all