> Markdown version of [/jobs/ext/1918106-senior-hpc-systems-engineer](https://www.wearedevelopers.com/jobs/ext/1918106-senior-hpc-systems-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior HPC Systems Engineer - **Company:** Parallel Works Inc. - **Location:** Chicago, IL, United States (Remote available) - **Experience:** Expert - **Salary:** $115,000.0 - $140,000.0 - **Contract:** Permanent contract - **Skills:** Computing Platforms, Systems Engineering, Bash Shell, Ubuntu (Operating System), Cloud Computing, Computer Clusters, Nvidia CUDA, Debian Linux, Linux, DevOps, Federal Information Processing Standards (FIPS), Firmware, InfiniBand, Python (Programming Language), OpenShift, Prolog, Red Hat Enterprise Linux, Ansible, Prometheus, Security Content Automation Protocol, Scripting, Grafana, SC Clearance, Kubernetes, Bare Metal, Slurm, 3-tier Architectures, Terraform - **Published:** August 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=15cef5e514b5208d ## About the Role * 10 or more years operating production Linux systems across more than one distribution family. RHEL, Rocky, or Alma on the Government side and Debian or Ubuntu on the commercial GPU side, since those clusters usually ship Ubuntu. Kernel and network tuning, systemd, cgroups, NUMA. * Production Slurm administration. You have configured, debugged, and upgraded a scheduler other people depended on. * At least one parallel or high throughput filesystem in production, plus InfiniBand or RoCE fabric operations. * NVIDIA GPU node operations at multi-node scale, including driver stack management and fault triage. * Experience on customer owned or on-premises clusters as well as public cloud, including bare metal provisioning and out of band management. * DevOps experience with infrastructure as code frameworks such as Ansible or Terraform. * Proficiency with common scripting languages such as Bash and Python. * United States citizenship and eligibility for a Secret clearance, since the work reaches export controlled Government environments. An active clearance helps. We sponsor candidates who are eligible but not currently cleared. You do not need every item on this list. If you have most of it and work well with other people, apply. Preferred Qualifications * Time at a Government supercomputing center, national laboratory, or university research computing center. * Work with STIG, SCAP, Tenable, eMASS, or RMF, or time inside FedRAMP or Impact Level boundaries. * Running GPU workloads on Kubernetes or OpenShift with Helm and operators, and using HPC containers such as Apptainer, Enroot, or Pyxis. * Running PBS Pro alongside Slurm, and monitoring with Prometheus and Grafana, including utilization and chargeback reporting. ## Description Parallel Works is hiring a Senior HPC Systems Engineer to build and run the clusters behind our defense and research programs. The work covers GPU node bring-up, Slurm configuration, fabric and storage troubleshooting, security hardening, and Tier 3 escalation. The computing environments are hybrid. Some clusters are customer owned hardware on site, some run in accredited Government cloud regions, and some are dedicated GPU clusters at commercial providers. On several programs the on-premises systems carry the primary load and cloud takes the overflow. The position is senior: it handles the escalations the rest of the team cannot resolve, and it trains the junior engineers. What you will do * Cluster operations: build and operate production Slurm clusters. slurmctld and slurmdbd, partitions and QOS, accounts and fair share, GPU GRES, prolog and epilog, cgroup enforcement. * Hybrid federation: connect customer owned clusters to the control plane, reconciling their site scheduler, storage, and identity source so accounts and allocations behave the same in every venue. * On-premises hardware: bare metal provisioning, out of band management, firmware, rack networking, and fault coordination with site staff or vendors. * GPU and fabric: validate GPU nodes before users arrive. Driver and CUDA stack, DCGM health checks, XID triage, fabric manager and NVLink checks, InfiniBand verification, NCCL tuning. * Storage and automation: tune parallel and high throughput storage, and write the Ansible, Terraform, and image build pipelines that make a cluster reproducible. * Security and escalation: STIG hardening, scan remediation, FIPS validated cryptography, security package artifacts, Tier 3 escalations, and a share of the on call rotation. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Discover the open source trio you didn’t expect: .NET and PostgreSQL on Linux](https://www.wearedevelopers.com/videos/2042-discover-the-open-source-trio-you-didn-t-expect-net-and-postgresql-on-linux) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 134 - Where pixels sing?](https://www.wearedevelopers.com/magazine/477-dev-digest-134-where-pixels-sing) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)