Hpc And It Infrastructure Engineer

Nostrum Biodiscovery
Madrid, Spain
3 days ago
Apply on www.buscojobs.com.es
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Amazon Web Services Bash Shell Configuration Management Compilers Cyber Security Nvidia CUDA Monitoring of Systems Identity and Access Management Issue Tracking Systems Python (Programming Language) Linux System Administration Network Service
+9 more
Open Source Technology Performance Tuning Security Information and Event Management Software Vulnerability Management Cloud Platform System Information Technology Slurm Hardware Infrastructure User Accounts

Job description

We are seeking an HPC and IT Infrastructure Engineer to manage the computing environment our science runs on.In this role, you will be in charge of the clusters where our simulations execute, the network and devices our team works from, and the security controls that protect both.You will work directly with computational chemists and software engineers whose results depend on the cluster being fast and available, and you will have a real say in how the environment is designed.Expect roughly 60% of your time on HPC and infrastructure engineering and 40% on IT services and security operations.Target start date: immediately.Job ResponsibilitiesAdminister the on-premises HPC cluster and AWS ParallelCluster environments, including scheduler, queues, accounting, and capacity, with an eye on cloud spend.Maintain the scientific software stack (modules, containers, compilers, MPI, CUDA) and troubleshoot performance and reliability issues alongside the researchers running the jobs.Support clusters dedicated to client projects, including maintenance and documentation.Operate the office network, servers, storage, and backups.Administer internally hosted services and applications, including the intranet and the ticketing system.Provide first- and second-line user support through our internal ticketing system, covering user onboarding and offboarding, credentials and access control, and the management of laptops, workstations, and company devices.Manage endpoint security, alert response, and vulnerability remediation, and maintain the technical controls required by ISO/IEC **.Automate recurring operational work and keep configuration standards documented.Required SkillsSolid Linux systems administration experience, including networking, storage, and troubleshooting at the OS level.Hands-on experience administering HPC clusters with Slurm: partitions, QOS, accounting, and job troubleshooting.Experience building or maintaining scientific software environments: module systems, containers, compilers, and MPI.Experience with AWS, particularly compute, storage, and networking services, and cost awareness in cloud environments.Scripting and automation (Bash, Python) and configuration management or infrastructure as code.Experience managing endpoints and user accounts across mixed operating systems, including identity and access management.Working knowledge of information security practice: access control, endpoint protection, vulnerability management, and incident response.Ability to document your work clearly and to communicate with non-specialist users.We will value experience inAWS Parallel Cluster or other cloud HPC deployments.GPU infrastructure: CUDA toolchains, drivers, and scheduling GPU workloads.Compiling and installing open-source or scientific software, resolving dependency issues, and delivering functional builds to end users.ISO/IEC ** implementation or audit support and GDPR technical controls.Security monitoring and SIEM tooling, and infrastructure monitoring.Working in a client-facing or regulated environment where infrastructure is part of a service commitment.Benefits of working at NostrumCompetitive salary based on experience and Barcelona market benchmarks, plus an annual bonus tied to company and department/individual performance Flexible working hours and teleworking policy.Health care insurance.Possibility to have food, transportation, or nursery tickets (tax benefits).As part of the career plan and development plans, Nostrum will facilitate all necessary training and future certifications to bring this role to the next level.Exposure to international projects and teams, with offices in Barcelona and Boston and a growing footprint in Asia#J-*****-Ljbffr

Requirements

Solid Linux systems administration experience, including networking, storage, and troubleshooting at the OS level. Hands-on experience administering HPC clusters with Slurm: partitions, QOS, accounting, and job troubleshooting. Experience building or maintaining scientific software environments: module systems, containers, compilers, and MPI. Experience with AWS, particularly compute, storage, and networking services, and cost awareness in cloud environments. Scripting and automation (Bash, Python) and configuration management or infrastructure as code. Experience managing endpoints and user accounts across mixed operating systems, including identity and access management. Working knowledge of information security practice: access control, endpoint protection, vulnerability management, and incident response. Ability to document your work clearly and to communicate with non-specialist users. We will value experience in AWS Parallel Cluster or other cloud HPC deployments. GPU infrastructure: CUDA toolchains, drivers, and scheduling GPU workloads. Compiling and installing open-source or scientific software, resolving dependency issues, and delivering functional builds to end users. ISO/IEC ***** implementation or audit support and GDPR technical controls. Security monitoring and SIEM tooling, and infrastructure monitoring. Working in a client-facing or regulated environment where infrastructure is part of a service commitment. Benefits of working at Nostrum

Benefits & conditions

Competitive salary based on experience and Barcelona market benchmarks, plus an annual bonus tied to company and department/individual performance Flexible working hours and teleworking policy. Health care insurance. Possibility to have food, transportation, or nursery tickets (tax benefits). As part of the career plan and development plans, Nostrum will facilitate all necessary training and future certifications to bring this role to the next level. Exposure to international projects and teams, with offices in Barcelona and Boston and a growing footprint in Asia #J-*****-Ljbffr

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.buscojobs.com.es
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:51 min

Evolution of custom compilers and virtual machines

Florian Rappl · LIVE

2:12 min

Integrating tool definitions for local bash shell execution

Michał Michalczuk Michał Michalczuk · Europe 2026 Virtual

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all