HPC Consultant

CYNET SYSTEMS INC.
Fremont, CA, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$108,160.0 - $118,560.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Application Performance Management Microsoft Azure Bash Shell Compilers Nvidia CUDA Software Debugging File Systems General-Purpose Computing on Graphics Processing Units Job Scheduling Python (Programming Language) Network Layer
+14 more
Linux System Administration OpenMP Performance Tuning Runbook Software Project Management Systems Architecture Toolchain Scripting Google Cloud High Performance Computing Performance Testing Storage Technologies Information Technology Slurm

Job description

  • Responsible for designing, optimizing, and supporting high-performance computing (HPC) environments including cluster scheduling, storage performance, application optimization, and system tuning.
  • The role involves improving workload efficiency, supporting HPC applications, and ensuring optimal performance across compute, storage, and network layers in large-scale production environments., * Design, configure, tune, and optimize SLURM partitions, queues, QoS, and scheduling policies.
  • Analyze job scheduling behavior, bottlenecks, and resource contention issues.
  • Troubleshoot job failures and performance degradation in HPC environments.
  • Implement scheduling policies such as fair-share, backfill, and reservations.
  • Lead HPC storage benchmarking and performance validation activities.
  • Analyze HPC workload I/O patterns and recommend storage architectures.
  • Support storage procurement decisions including performance and sizing analysis.
  • Collaborate with vendors and internal teams during proof-of-concept evaluations.
  • Build, configure, and maintain HPC applications, compilers, and software stacks.
  • Optimize application performance using MPI, OpenMP, and GPU acceleration where applicable.
  • Manage environment modules and software management frameworks.
  • Perform system-level tuning across compute, memory, network, and storage systems.
  • Diagnose and resolve node-level issues involving CPU, GPU, interconnects, and OS configurations.
  • Create runbooks, performance baselines, and troubleshooting documentation.
  • Support cluster upgrades, expansions, and infrastructure lifecycle activities.
  • Collaborate with researchers, application owners, and infrastructure teams.
  • Translate workload requirements into optimized HPC configurations.
  • Provide technical guidance and recommendations to stakeholders and leadership.

Requirements

  • Eight to twelve years of hands-on HPC engineering experience in production environments.
  • Strong expertise in SLURM configuration, tuning, and troubleshooting.
  • Strong knowledge of Linux operating systems.
  • Experience with HPC storage systems and I/O performance analysis.
  • Experience building, installing, and optimizing HPC applications and scientific software stacks.
  • Experience with MPI, OpenMP, and HPC toolchains.
  • Strong scripting skills in Bash and Python.
  • Experience with performance analysis and debugging tools.
  • Strong understanding of HPC system architecture and workload optimization.

Experience:

  • Experience designing and tuning HPC cluster scheduling policies including fair-share, backfill, and reservations.
  • Experience in HPC storage benchmarking using tools such as IOR, FIO, MDTest, and IOzone.
  • Experience analyzing I/O patterns and mapping workloads to storage architectures.
  • Experience supporting application optimization using compilers and libraries.
  • Experience in system-level performance tuning across compute, storage, and network layers.
  • Experience supporting cluster upgrades, expansions, and hardware refresh activities., * Experience with GPU-based HPC workloads (CUDA, ROCm).
  • Exposure to cloud HPC environments (Azure, AWS, Google Cloud Platform).
  • Experience with parallel file systems such as Lustre or IBM Spectrum Scale.
  • Experience working with vendors for HPC hardware and storage evaluations.

Skills:

  • SLURM scheduling and cluster management.
  • Linux system administration.
  • HPC storage and I/O performance tuning.
  • MPI and OpenMP programming models.
  • HPC compilers and toolchains (GCC, Intel, NVIDIA HPC SDK).
  • Performance analysis tools.
  • Python and Bash scripting.
  • Environment modules (Lmod).
  • HPC system architecture and optimization.
  • GPU computing (preferred).

Qualification And Education:

  • Bachelor s or Master s degree in Computer Science, Engineering, or related field preferred.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

2:50 min

Introduction and the value of runbooks

Hila Fish · WWC 2023

4:37 min

Simplifying parallel programming with the CUDA ecosystem

Paul Graham Paul Graham · LIVE

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · WWC Europe 2026

2:04 min

Insights on transitioning from supercomputing to technical education

Andrew Holway · LIVE

1:32 min

Structuring automated incident workflows between runbooks and raw models

Aram Hakobyan Aram Hakobyan +1 · WWC Europe 2026

Videos

See all

Related articles

See all