TS/SCI HPC Systems Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+1 more
Job description
A federal IT services client of Insight Global is hiring for a highly skilled HPC Systems Engineer to join their team full time in Charlottesville, VA. This role requires an active TS/SCI clearance and is 5 days/week on site. Relocation packages are available!
The HPC (High Performance Computing) Systems Engineer will work directly with engineers, analysts, and researchers to support job execution, troubleshoot workload failures, and improve the performance and efficiency of compute workloads running on HPC clusters. The Engineer will assist users with scheduler job scripts, application execution, and workload performance troubleshooting while promoting HPC best practices for efficient cluster utilization. This role serves as the primary interface between mission users and HPC platform infrastructure teams.
Key Responsibilities:
-
Provide direct support to users running computational workloads on HPC clusters (classified & unclassified)
-
Assist with creating, submitting, and troubleshooting job scripts (Slurm, PBS), including CPU/GPU resource allocation
-
Diagnose and resolve failing, slow, or hanging jobs (including MPI, parallel, and GPU workloads)
-
Support application setup, compilation, and execution in Linux-based HPC environments
-
Advise users on best practices to improve job performance, efficiency, and resource utilization
-
Monitor workload usage and recommend optimizations to maximize cluster throughput
-
Develop and maintain automation scripts/tools (Bash/Python) and manage them in version control (Git)
-
Collaborate with infrastructure teams and maintain documentation to resolve system issues and support users
Requirements
Active TS/SCI clearance
-
ONE of the following certifications: Security+, CCNA Security, CySA+, GICSP, GSEC, CND, SSCP, CAP, CASP+, CISM, CISSP, GSLC, CCISO, HCISPP
-
5+ years of experience working in Linux environments supporting distributed compute workloads or HPC cluster platforms
-
Experience executing or troubleshooting workloads using HPC workload schedulers such as Slurm, PBS, Torque, or similar systems
-
Experience administering command-line Linux systems including scripting (Bash, python, etc.) and troubleshooting applications in multi-user server environments.
-
Experience supporting systems within DoD/DoW or IC environments
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on juju.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Dev Digest 134 - Where pixels sing?
Top 6 Hackathons for Developers in 2023
Highest Paying Tech Companies for Developers
Dev Digest 131 - AI'm not sure about OSS