nHPC Systems Engineer

SBS SYSTEMS LLC
Chantilly, VA, United States
23 days ago
Apply on jobs.military.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Amazon Web Services Systems Engineering Bash Shell Configuration Management Distributed Systems Job Scheduling Python (Programming Language) Performance Tuning Scripting High Performance Computing Containerization Slurm
+1 more
Docker

Requirements

The ideal candidate will have strong experience with large-scale, on-premises or hybrid HPC infrastructures and a passion for optimizing system performance. While the primary focus is on traditional HPC environments, familiarity with \nAWS-based technologies is beneficial for supporting hybrid or cloud-extended HPC use cases.\n, * Demonstrated experience supporting High Performance Computing (HPC) or large-scale distributed computing environments.\n

  • Strong proficiency in Linux/Unix system administration.\n
  • Experience with job schedulers such as Slurm, PBS, or LSF.\n
  • Understanding of parallel and distributed computing concepts.\n
  • Familiarity with high-performance networking and parallel storage systems.\n
  • Scripting experience using Bash, Python, or similar languages.\n
  • Strong analytical and problem-solving skills.\n
  • Excellent communication and collaboration abilities.\n, * Experience supporting federal or defense customers.\n
  • Familiarity with GPU-accelerated computing and performance optimization.\n
  • Experience with container technologies such as Singularity/Apptainer or Docker.\n
  • Exposure to configuration management or automation tools.\n
  • Experience with AWS-based technologies (e.g., AWS services supporting HPC or hybrid environments).\n
  • Relevant industry or technical certifications.\n

Benefits & conditions

  • Manage and support job scheduling and workload management platforms (e.g., Slurm, PBS, or similar).\n
  • Monitor and tune system performance to ensure efficient utilization of compute, storage, and network resources.\n
  • Support high-speed interconnects (e.g., InfiniBand) and parallel file systems such as Lustre, GPFS, or BeeGFS.\n
  • Collaborate with engineers, scientists, and mission stakeholders to translate computational requirements into effective technical solutions.\n
  • Implement automation using scripting languages such as Bash or Python.\n
  • Ensure system reliability, security, and compliance with organizational and government standards.\n
  • Provide technical documentation and user support.\n
  • Contribute to capacity planning and infrastructure enhancements.\n
  • Nice to Have: Support the integration or extension of HPC workloads using AWS-based technologies for hybrid computing scenarios.\n

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.military.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

1:04 min

Introduction to Bitcoin script parsing tools

Steve Shadders · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:53 min

Evaluating traditional scripting languages for modern development tasks

Jens Knipper Jens Knipper · Europe 2026 Virtual

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all