Senior HPC Engineer

Do IT Now
München, Germany
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Languages
English, German
Job source

Tech stack

Artificial Intelligence Software Design Documents File Systems General Parallel File Systems InfiniBand Red Hat Enterprise Linux AI Infrastructure SAP Ariba Slurm Docker Confluent

Job description

Do IT Now is looking for a Senior HPC Engineer to join an established technical group delivering Worldwide HPC and AI infrastructure services to customers across research, automotive, life sciences, and manufacturing.

The required role is responsible for design, deployment and operational excellence of HPC and AI clusters across customer sites and managed environments.

Cluster Design and Deployment: Lead the architectural design and on-site or remote deployments of HPC and AI clusters. Covering GPU systems, high-speed interconnects, parallel file systems and management infrastructure.

Workload Manager: Configure, tune and maintain schedulers like Slurm, PBS and Gridengine.

Fabric Engineering: Deploy and operate high-speed network fabrics (InfiniBand NDR, RoCE), including topology design, SM configuration, validation and testing.

Customer Interface: Serve as technical counterpart for customer or partner architects and operation leads.

Technical Writing and Training: Author runbooks, design documents, acceptance documents and handover materials. Provide technical training sessions for customers and mentor junior engineers.

Possibility to work remotely from all over Germany but with the requirement to complete the training period on site and have meetings at the Munich workplace.

Requirements

Essential skills

  • Over 5 years of experience in similar roles.
  • Excellent knowledge of English and German (at least B2).
  • Experience in enterprise Linux environments.
  • Production experience with multiple Workload Managers, like Slurm, PBS or Gridengine.
  • Hands-on experience with InfiniBand and comparable high-speed interconnects.
  • Practical experience with at least one parallel file system (BeeGFS, Lustre, Storage Scale/GPFS).
  • Punctual and constructive participation in internal and client meetings.
  • Excellent oral and written communication skills in technical contexts.

Preferential requirements.

  • Experience with European HPC customers, like EuroHPC, automotive CAE or life sciences.
  • Experience with cluster managers (WareWulf, Confluent, SMC xSCale, BCM, OpenCHAMI).
  • Experience with HA tooling for service nodes (Pacemaker, Corosync, fencing).
  • Familiarity with container runtimes (Apptainer, Podman, Enroot, Pyxis)
  • Commercial awareness and good understanding of the HPC and AI market, * Proactive attitude and results-oriented mindset.
  • Ability to innovate.
  • Problem-solving skills.
  • Excellent interpersonal and teamwork skills.
  • Strong operational and organizational skills in technologically advanced work environments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on de.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · WWC Europe 2026

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:47 min

Exploring career opportunities and recruitment open positions

Kurt Eder · LIVE

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · WWC Europe 2026

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all