HPC consultant

RPA Infotech Inc.
Fremont, CA, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Data Analysis Application Performance Management Microsoft Azure Bash Shell CentOS Cloud Computing Compilers Nvidia CUDA Linux File Systems Ethernet
+17 more
InfiniBand Job Scheduling OpenMP Package Management Systems Performance Tuning Red Hat Enterprise Linux Software Project Management Toolchain Scripting Google Cloud High Performance Computing Build Management Perf (Linux) Storage Technologies Performance Monitor Slurm Nvme

Job description

  • Design, configure, tune, and optimize SLURM partitions, queues, QoS, and scheduling policies to maximize cluster utilization and workload efficiency.
  • Perform in-depth analysis of job scheduling behavior, bottlenecks, and resource contention.
  • Troubleshoot job failures, performance degradation, and scheduler-related issues in production HPC environments.
  • Implement fair-share, backfill, reservations, and policy-driven scheduling as required.

Storage Benchmarking & Procurement Support

  • Lead HPC storage performance benchmarking using industry-standard tools (e.g., IOR, FIO, MDTest, IOzone).
  • Analyze I/O patterns of HPC workloads and map them to appropriate storage architectures (parallel file systems, NVMe, Lustre, Spectrum Scale, etc.).
  • Provide technical input for storage selection and procurement, including performance expectations, sizing, and cost-performance tradeoffs.
  • Collaborate with vendors and internal teams during POCs and performance validation exercises.

HPC Application Build & Optimization

  • Build, install, configure, and maintain HPC applications, compilers, libraries, and scientific software stacks.
  • Optimize application performance using MPI, OpenMP, GPU acceleration (where applicable), and tuned math libraries.
  • Support multiple compiler toolchains (GCC, Intel, LLVM, NVIDIA HPC SDK, etc.).
  • Implement and manage environment modules (Lmod) or similar software management frameworks.

System Performance & Operations

  • Conduct system-level performance tuning across compute, memory, network, and storage layers.
  • Diagnose node-level issues involving CPU, GPU, interconnects (InfiniBand/Ethernet), and OS configurations.
  • Create operational runbooks, performance baselines, and troubleshooting documentation.
  • Support cluster upgrades, expansions, and hardware refresh activities.

Collaboration & Delivery

  • Work closely with application owners, researchers, and infrastructure teams to meet aggressive delivery timelines.
  • Translate workload requirements into practical HPC configurations and optimizations.
  • Provide clear technical guidance and recommendations to leadership and stakeholders.

Requirements

Core HPC Skills

  • 8-12+ years of hands-on HPC engineering experience in production environments.
  • Strong expertise with SLURM (configuration, tuning, troubleshooting).
  • Solid understanding of Linux systems (RHEL/CentOS/Rocky/Alma preferred).
  • Deep knowledge of HPC storage systems and I/O performance analysis.
  • Proven experience building and optimizing HPC applications and libraries.

Technical Proficiency

  • MPI implementations (Open MPI, MPICH), OpenMP
  • Compilers and toolchains (GCC, Intel, NVIDIA HPC SDK)
  • Performance tools (perf, vtune, nvprof/nsys, IB diagnostics)
  • Environment modules (Lmod), package managers (Spack preferred)
  • Bash/Python scripting for automation and diagnostics

Nice to Have

  • Experience with GPU-based HPC workloads (NVIDIA CUDA, ROCm).
  • Exposure to cloud-based HPC (Azure, AWS, Google Cloud Platform).
  • Familiarity with parallel file systems such as Lustre or IBM Spectrum Scale.
  • Vendor engagement experience for HPC hardware/storage evaluations.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · WWC Europe 2026

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · WWC Europe 2026

1:31 min

Profiling computing workloads across environments using Performance Studio

Andrew Wafaa Andrew Wafaa · WWC 2024

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

Videos

See all

Related articles

See all