HPC Consultant
CYNET SYSTEMS INC.
Fremont, CA, United States
3 months ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$108,160.0 - $118,560.0
Working hours
Regular working hours
Job source
Tech stack
Amazon Web Services
Application Performance Management
Microsoft Azure
Bash Shell
Compilers
Nvidia CUDA
Software Debugging
File Systems
General-Purpose Computing on Graphics Processing Units
Job Scheduling
Python (Programming Language)
Network Layer
+14 more
Linux System Administration
OpenMP
Performance Tuning
Runbook
Software Project Management
Systems Architecture
Toolchain
Scripting
Google Cloud
High Performance Computing
Performance Testing
Storage Technologies
Information Technology
Slurm
Job description
- Responsible for designing, optimizing, and supporting high-performance computing (HPC) environments including cluster scheduling, storage performance, application optimization, and system tuning.
- The role involves improving workload efficiency, supporting HPC applications, and ensuring optimal performance across compute, storage, and network layers in large-scale production environments., * Design, configure, tune, and optimize SLURM partitions, queues, QoS, and scheduling policies.
- Analyze job scheduling behavior, bottlenecks, and resource contention issues.
- Troubleshoot job failures and performance degradation in HPC environments.
- Implement scheduling policies such as fair-share, backfill, and reservations.
- Lead HPC storage benchmarking and performance validation activities.
- Analyze HPC workload I/O patterns and recommend storage architectures.
- Support storage procurement decisions including performance and sizing analysis.
- Collaborate with vendors and internal teams during proof-of-concept evaluations.
- Build, configure, and maintain HPC applications, compilers, and software stacks.
- Optimize application performance using MPI, OpenMP, and GPU acceleration where applicable.
- Manage environment modules and software management frameworks.
- Perform system-level tuning across compute, memory, network, and storage systems.
- Diagnose and resolve node-level issues involving CPU, GPU, interconnects, and OS configurations.
- Create runbooks, performance baselines, and troubleshooting documentation.
- Support cluster upgrades, expansions, and infrastructure lifecycle activities.
- Collaborate with researchers, application owners, and infrastructure teams.
- Translate workload requirements into optimized HPC configurations.
- Provide technical guidance and recommendations to stakeholders and leadership.
Requirements
- Eight to twelve years of hands-on HPC engineering experience in production environments.
- Strong expertise in SLURM configuration, tuning, and troubleshooting.
- Strong knowledge of Linux operating systems.
- Experience with HPC storage systems and I/O performance analysis.
- Experience building, installing, and optimizing HPC applications and scientific software stacks.
- Experience with MPI, OpenMP, and HPC toolchains.
- Strong scripting skills in Bash and Python.
- Experience with performance analysis and debugging tools.
- Strong understanding of HPC system architecture and workload optimization.
Experience:
- Experience designing and tuning HPC cluster scheduling policies including fair-share, backfill, and reservations.
- Experience in HPC storage benchmarking using tools such as IOR, FIO, MDTest, and IOzone.
- Experience analyzing I/O patterns and mapping workloads to storage architectures.
- Experience supporting application optimization using compilers and libraries.
- Experience in system-level performance tuning across compute, storage, and network layers.
- Experience supporting cluster upgrades, expansions, and hardware refresh activities., * Experience with GPU-based HPC workloads (CUDA, ROCm).
- Exposure to cloud HPC environments (Azure, AWS, Google Cloud Platform).
- Experience with parallel file systems such as Lustre or IBM Spectrum Scale.
- Experience working with vendors for HPC hardware and storage evaluations.
Skills:
- SLURM scheduling and cluster management.
- Linux system administration.
- HPC storage and I/O performance tuning.
- MPI and OpenMP programming models.
- HPC compilers and toolchains (GCC, Intel, NVIDIA HPC SDK).
- Performance analysis tools.
- Python and Bash scripting.
- Environment modules (Lmod).
- HPC system architecture and optimization.
- GPU computing (preferred).
Qualification And Education:
- Bachelor s or Master s degree in Computer Science, Engineering, or related field preferred.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on dice.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
about 3 years ago
LM
Luis Minvielle
7 Cloud Computing Trends Coming in 2025 for Developers
over 2 years ago
AJ
Austin Joy
What Are The Top Skills Required For Azure Developers?
over 4 years ago
DD
Dilek Demir
Data Science & more: The Lopez dilemma
almost 6 years ago
JF
Jonas Fritzsch
Résumé-Driven Development: How IT trends affect the job market for software developers
almost 5 years ago
KD
Krissy Davis
Best Coding Boot Camps in Germany
about 3 years ago