HPC Systems Engineer - Security Architecture & Cluster Administration

Cyber Valley GmbH
Tübingen, Germany
6 days ago
Apply on de.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Shift work
Languages
English, German
Job source

Tech stack

Proxmox Kubernetes Security Microsoft Access Systems Engineering Bash Shell Ubuntu (Operating System) Cloud Computing Configuration Management Cyber Security Data Centers Linux File Systems
+21 more
Ethernet Identity and Access Management InfiniBand Intrusion Detection and Prevention Job Scheduling Python (Programming Language) Network Security Machine Learning Open Access Red Hat Enterprise Linux Ansible Software Vulnerability Management Weka Ceph (Software) Data Logging Data Processing High Performance Computing System Availability Information Technology Slurm Docker

Job description

Traditionally, HPC environments prioritized only performance and open access, with security as an afterthought. As their clusters grow and handle increasingly sensitive research data, security becomes mission-critical. They are looking for an HPC System Engineer who combines hands-on cluster operations with a security mindset. You will actively shape the security architecture of a live, production HPC / ML environment, keep their systems hardened, and build the processes and infrastructure that protect researchers, their data and their compute against a fast-evolving threat landscape., * Design and operate their HPC clusters across four data centers, including scheduler (SLURM), parallel filesystems, networks and accelerators, ensuring high availability and throughput for research workloads

  • Conceive and establish the security architecture of the Machine Learning Science Cloud and harden the HPC environment
  • Evolve the automated provisioning and configuration of heterogeneous compute, storage and network nodes (e.g. image-based provisioning, node lifecycle)
  • Run patch and vulnerability management - risk assessment across heterogeneous systems
  • Build and operate logging, monitoring and intrusion detection, and integrate HPC telemetry into both operational dashboards and incident-response workflows
  • Lead incident response for the clusters: detection, containment, forensic support and post-incident review
  • Automate operations and security policy as code (Ansible/IaC)
  • Advice researchers on efficient, secure cluster usage (job scheduling, data handling, access workflows) and derive requirements for our further roadmap from their scientific workloads

Requirements

  • Masters degree in Computer Science or a related field
  • In-depth IT security knowledge: system hardening, network security, applied cryptography and IAM - with the ability to derive architectural decisions from a threat model, not only to apply given baselines
  • Solid Linux experience in production environments (RHEL/Almalinux/Ubuntu)
  • Hands-on HPC background: Slurm, parallel file systems (Weka, Lustre, Ceph), GPU workloads and high-speed networks (InfiniBand, 400G Ethernet)
  • Experience with virtualization for management-plane and infrastructure services (Proxmox)
  • Strong scripting and automation skills (Bash, Python, Ansible) and experience with configuration management / Infrastructure-as-Code
  • A plus: security frameworks (ISO 27001, BSI Grundschutz), container security (Apptainer/Singularity/Docker) or offensive-security fundamentals
  • Independent, structured working style and good communication in English; German is a plus
  • A collaborative, user-facing mindset - comfortable supporting and advising researchers and translating their needs into platform design

Benefits & conditions

Pulled from the full job description

  • Flexible schedule, * Technically deep, architecturally open work that directly enables cutting-edge machine-learning research
  • Flexible working hours and the option to work partially from home
  • Working in an English-speaking, international team of HPC experts
  • A small, senior team with flat hierarchy where responsibility is split by domain
  • Ownership of a technical domain in a production environment of real scale
  • Professional development, conference attendance and real influence on our roadmap

About the company

The Cluster of Excellence “Machine Learning - New Perspectives for Science” together with the Tübingen AI Center at the University of Tübingen offers a position as, The ML Cloud team designs, runs and safeguards high-performance computing infrastructure tailored to machine learning workloads. Their state of the art systems are distributed across four data centers for the Tübingen AI Center, the Cluster of Excellence ‘Machine Learning: New Perspectives for Science’ and the Hertie Institute for AI in Brain Health. Every day, researchers unleash thousands of compute jobs on their systems, training frontier-scale neural networks and running experiments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on de.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:47 min

Exploring career opportunities and recruitment open positions

Kurt Eder · LIVE

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all