> Markdown version of [/jobs/ext/37290-hpc-consultant](https://www.wearedevelopers.com/jobs/ext/37290-hpc-consultant). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # HPC Consultant - **Company:** CYNET SYSTEMS INC. - **Location:** Fremont, CA, United States - **Experience:** Expert - **Salary:** $108,160.0 - $118,560.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Application Performance Management, Microsoft Azure, Bash Shell, Compilers, Nvidia CUDA, Software Debugging, File Systems, General-Purpose Computing on Graphics Processing Units, Job Scheduling, Python (Programming Language), Network Layer, Linux System Administration, OpenMP, Performance Tuning, Runbook, Software Project Management, Systems Architecture, Toolchain, Scripting, Google Cloud, High Performance Computing, Performance Testing, Storage Technologies, Information Technology, Slurm - **Published:** May 12, 2026 - **Apply:** https://www.dice.com/job-detail/7aac9a75-0585-4c01-9170-30eaa1c85e2b ## About the Role * Eight to twelve years of hands-on HPC engineering experience in production environments. * Strong expertise in SLURM configuration, tuning, and troubleshooting. * Strong knowledge of Linux operating systems. * Experience with HPC storage systems and I/O performance analysis. * Experience building, installing, and optimizing HPC applications and scientific software stacks. * Experience with MPI, OpenMP, and HPC toolchains. * Strong scripting skills in Bash and Python. * Experience with performance analysis and debugging tools. * Strong understanding of HPC system architecture and workload optimization. Experience: * Experience designing and tuning HPC cluster scheduling policies including fair-share, backfill, and reservations. * Experience in HPC storage benchmarking using tools such as IOR, FIO, MDTest, and IOzone. * Experience analyzing I/O patterns and mapping workloads to storage architectures. * Experience supporting application optimization using compilers and libraries. * Experience in system-level performance tuning across compute, storage, and network layers. * Experience supporting cluster upgrades, expansions, and hardware refresh activities., * Experience with GPU-based HPC workloads (CUDA, ROCm). * Exposure to cloud HPC environments (Azure, AWS, Google Cloud Platform). * Experience with parallel file systems such as Lustre or IBM Spectrum Scale. * Experience working with vendors for HPC hardware and storage evaluations. Skills: * SLURM scheduling and cluster management. * Linux system administration. * HPC storage and I/O performance tuning. * MPI and OpenMP programming models. * HPC compilers and toolchains (GCC, Intel, NVIDIA HPC SDK). * Performance analysis tools. * Python and Bash scripting. * Environment modules (Lmod). * HPC system architecture and optimization. * GPU computing (preferred). Qualification And Education: * Bachelor s or Master s degree in Computer Science, Engineering, or related field preferred. ## Description * Responsible for designing, optimizing, and supporting high-performance computing (HPC) environments including cluster scheduling, storage performance, application optimization, and system tuning. * The role involves improving workload efficiency, supporting HPC applications, and ensuring optimal performance across compute, storage, and network layers in large-scale production environments., * Design, configure, tune, and optimize SLURM partitions, queues, QoS, and scheduling policies. * Analyze job scheduling behavior, bottlenecks, and resource contention issues. * Troubleshoot job failures and performance degradation in HPC environments. * Implement scheduling policies such as fair-share, backfill, and reservations. * Lead HPC storage benchmarking and performance validation activities. * Analyze HPC workload I/O patterns and recommend storage architectures. * Support storage procurement decisions including performance and sizing analysis. * Collaborate with vendors and internal teams during proof-of-concept evaluations. * Build, configure, and maintain HPC applications, compilers, and software stacks. * Optimize application performance using MPI, OpenMP, and GPU acceleration where applicable. * Manage environment modules and software management frameworks. * Perform system-level tuning across compute, memory, network, and storage systems. * Diagnose and resolve node-level issues involving CPU, GPU, interconnects, and OS configurations. * Create runbooks, performance baselines, and troubleshooting documentation. * Support cluster upgrades, expansions, and infrastructure lifecycle activities. * Collaborate with researchers, application owners, and infrastructure teams. * Translate workload requirements into optimized HPC configurations. * Provide technical guidance and recommendations to stakeholders and leadership. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [CUDA Python: GPU programming for the modern developer](https://www.wearedevelopers.com/videos/100221-cuda-python-gpu-programming-for-the-modern-developer) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/1521-accelerating-python-on-gpus) - [3 Key Steps for Optimizing DevOps Workflows](https://www.wearedevelopers.com/videos/962-3-key-steps-for-optimizing-devops-workflows) - [Beyond Chat: AI Workflows That Actually Investigate Alerts (So You Don't Have To Know Everything)](https://www.wearedevelopers.com/videos/100308-beyond-chat-ai-workflows-that-actually-investigate-alerts-so-you-don-t-have-to-know-everything) ## Related Articles - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Data Science & more: The Lopez dilemma](https://www.wearedevelopers.com/magazine/10-data-science-more-the-lopez-dilemma) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Best Coding Boot Camps in Germany](https://www.wearedevelopers.com/magazine/237-best-coding-boot-camps-in-germany)