> Markdown version of [/jobs/ext/2522898-high-performance-computing-hpc-sa2-government](https://www.wearedevelopers.com/jobs/ext/2522898-high-performance-computing-hpc-sa2-government). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # High-Performance Computing (HPC) (SA2) (Government) - **Company:** AT&T Inc. - **Location:** Columbia, MD, United States - **Contract:** Permanent contract - **Skills:** Microsoft Windows, Confluence, JIRA, Bash Shell, Linux, File Systems, Fault Tolerance, General Parallel File Systems, InfiniBand, Nagios, Performance Tuning, Ansible, Prometheus, Transmission Control Protocol (TCP), High Performance Computing, System Availability, Grafana, Git, Infrastructure Automation Frameworks, Slurm - **Published:** August 7, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/high-performance-computing-hpc-sa2-government-columbia-md-usa-58857588 ## About the Role FULL_TIME management tools: Slurm, git, Salt, Ansible * Provide user support and escalate/status communication to agency management * Optimize operations via resource utilization and capacity planning * Harden, patch, and tune Linux/UNIX/Windows systems for reliability and performance * Provide detailed analysis for escalated tickets and support dispatch/hardware resolution Tasks * B.S. in a technical discipline and 5 years' experience as a System Administrator in similar scope or 10 years' experience in lieu of degree * DoD 8570 IAT II level certification required * TS/SCI with polygraph clearance * Experience with HPC clusters, Lustre/GPFS, InfiniBand/Slingshot (or similar interconnects) Key requirements * Medical/Dental/Vision coverage * 401(k) plan * Tuition reimbursement program * Paid Time Off and Holidays * Paid Parental Leave * Disability Benefits ## Description Experteer Overview In this HPC Systems Administrator role, you will sustain and operate Linux-based HPC clusters across two sites, ensuring high availability and performance for government IT workloads. You'll support transition of new capabilities into operations and collaborate with SRE teams under government policies. The role emphasizes fault-tolerant operations, incident response, and performance optimization. This is a field-facing position on a government site with a clear mission to enable secure, scalable HPC infrastructure. Compensation / Benefits * Operate and sustain Linux-based HPC clusters and parallel file systems across two sites * Monitor, troubleshoot, and perform routine maintenance; incident response * Install/configure Linux OS, file systems, and TCP/IP networking; fix OS/app issues * Automate and administer via Bash scripting; install software as needed * Utilize observability tools: Jira, Confluence, Grafana, Prometheus, Nagios * Support HPC workload and configuration management tools: Slurm, git, Salt, Ansible * Provide user support and escalate/status communication to agency management * Optimize operations via resource utilization and capacity planning * Harden, patch, and tune Linux/UNIX/Windows systems for reliability and performance * Provide detailed analysis for escalated tickets and support dispatch/hardware resolution Tasks * B.S. in a technical discipline and 5 years' experience as a System Administrator in similar scope or 10 years' experience in lieu of degree * DoD 8570 IAT II level certification required * TS/SCI with polygraph clearance * Experience with HPC clusters, Lustre/GPFS, InfiniBand/Slingshot (or similar interconnects) Key requirements * Medical/Dental/Vision coverage * 401(k) plan * Tuition reimbursement program * Paid Time Off and Holidays * Paid Parental Leave * Disability Benefits ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Collaboration Quantified: Lessons from Open Source Developer Networks](https://www.wearedevelopers.com/videos/1422-collaboration-quantified-lessons-from-open-source-developer-networks) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 134 - Where pixels sing?](https://www.wearedevelopers.com/magazine/477-dev-digest-134-where-pixels-sing) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)