> Markdown version of [/jobs/ext/3433785-hpc-systems-administrator](https://www.wearedevelopers.com/jobs/ext/3433785-hpc-systems-administrator). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # HPC Systems Administrator - **Company:** 5C DATA CENTERS USA INC. - **Location:** Springfield, OH, United States (Remote available) - **Experience:** Expert - **Salary:** $120,000.0 - $135,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computing Platforms, User Authentication, Intelligent Platform Management Interface, Bash Shell, BIOS, Command-Line Interface, Cloud Computing, Computer Clusters, Network Congestion, Data Centers, Linux, File Systems, Domain Name System (DNS), Ethernet, Network Interface Controllers, Firmware, General-Purpose Computing on Graphics Processing Units, InfiniBand, IP Addressing, Python (Programming Language), Linux System Administration, Networking Basics, Network Diagnostics, Routing, PCI Express, Windows PowerShell, Remote Direct Memory Access, Red Hat Enterprise Linux, Remote Administration, Ansible, System Software, System Testing, Virtual Local Area Networks, AI Infrastructure, Diagnostic Tools, Scripting, Graphics Processing Unit (GPU), Cloud Platform System, High Performance Computing, Reliability of Systems, Computerised Systems, Bare Metal, Hardware Asset Management, Hardware Infrastructure - **Published:** September 22, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=51027d150b45b9c8 ## About the Role * Approximately 5 years of experience in Linux systems administration, data center operations, infrastructure support, HPC, cloud infrastructure, or a related technical role. Equivalent hands-on experience, education, or technical training may also be considered. * Working knowledge of Linux operating systems and command-line administration. * Basic understanding of enterprise server hardware and common server components. * Experience troubleshooting hardware or operating-system issues using logs and diagnostic tools. * Familiarity with networking fundamentals including IP addressing, DNS, interfaces, routing, and basic connectivity troubleshooting. * Familiarity with server hardware concepts including CPU, memory, storage, PCIe devices, NICs, power supplies, and firmware. * Exposure to GPU infrastructure or an interest in developing expertise with NVIDIA accelerated computing platforms. * Familiarity with out-of-band management technologies such as IPMI, Redfish, iDRAC, iLO, or equivalent tools. * Basic scripting or automation experience using Bash, Python, PowerShell, Ansible, or similar technologies. * Ability to follow technical procedures and troubleshoot problems methodically. * Ability to recognize when an issue requires escalation and clearly communicate collected findings. * Strong written communication and documentation skills. * Ability to work collaboratively across technical teams. * Participate in an on-call rotation. ## Description As an HPC Systems Administrator, you will support the day-to-day operation, maintenance, and troubleshooting of HPC and AI infrastructure across data center and cloud environments. This role focuses primarily on the infrastructure and hardware layer, including Linux operating systems, GPU servers, high-speed networking, storage connectivity, firmware, drivers, and out-of-band management. You will work alongside senior Systems Administrators, Network Engineering, Data Center Operations, Deployment Engineering, vendors, and other technical teams to troubleshoot infrastructure issues, restore systems to service, perform routine maintenance, and improve the reliability of production HPC environments. This position is well suited for someone with a foundation in Linux systems administration or data center infrastructure who wants to develop deeper expertise in HPC, GPU computing, high-performance networking, and large-scale AI infrastructure. How We Work at 5C Our core values guide how we collaborate, make decisions, support one another, and serve our customers. We're looking for people who embrace them and help us build something great. What You Will Do HPC Infrastructure Operations * Support and maintain Linux-based HPC and AI computing environments, including large-scale NVIDIA GPU clusters, while troubleshooting hardware, operating system, networking, storage, driver, and firmware issues across bare-metal, virtualized, containerized, and cloud-hosted platforms. * Investigate and resolve infrastructure and compute-node issues, document technical findings, maintain operational runbooks and support procedures, and collaborate with senior engineers to escalate complex problems and improve overall system reliability. GPU and Accelerated Computing Systems * Assist with the installation, configuration, validation, monitoring, and troubleshooting of NVIDIA GPU infrastructure, including drivers, firmware, system software, and GPU health management tools, while diagnosing performance, communication, and hardware-related issues. * Perform post-maintenance system validation, collect and analyze diagnostic data for GPU and infrastructure faults, and coordinate hardware replacement and RMA activities for GPUs, system boards, NICs, power supplies, and other server components. Linux System Administration * Administer and support enterprise Linux environments, including system services, filesystems, networking, storage, authentication, permissions, SSH, DNS, and NTP, while ensuring compliance with established security and operating system standards. * Troubleshoot operating system, CPU, memory, device-discovery, and kernel-related issues, and collect diagnostic data to investigate system crashes, hardware failures, and performance-related incidents. Server Hardware, Firmware, and Out-of-Band Management * Troubleshoot hardware issues across enterprise server platforms, including memory, PCIe devices, GPUs, NICs, storage systems, power supplies, fans, system boards, and cabling, while supporting firmware updates for BIOS, BMC, GPUs, drives, and other infrastructure components. * Utilize out-of-band management tools to perform remote administration, system health monitoring, hardware inventory and log collection, and collaborate with Data Center technicians on hardware replacements, break-fix activities, cabling validation, and post-maintenance testing. High-Performance Networking * Assist with troubleshooting Ethernet, InfiniBand, and RoCE connectivity within HPC and AI environments, including support for NVIDIA/Mellanox network adapters and resolution of link, interface, MTU, packet loss, VLAN, RDMA, and fabric-related issues. * Collect and analyze network diagnostic data, escalate complex switching and fabric issues to Network Engineering teams, and validate network performance following infrastructure changes, firmware upgrades, cable replacements, and new system deployments. Incident, Problem, and Change Management * Participate in incident response activities for production infrastructure environments, troubleshoot technical issues, provide status updates, and escalate incidents in accordance with established support procedures and service-level agreements. * Document troubleshooting activities, findings, corrective actions, and resolutions while following change-management processes, maintenance procedures, rollback plans, and on-call support requirements. Collaboration and Development * Collaborate with senior Systems Administrators and engineers to troubleshoot HPC infrastructure issues, follow established operational procedures, and participate in technical reviews, troubleshooting sessions, and post-incident analyses. * Contribute to runbooks, knowledge-base articles, and support documentation, share technical findings and lessons learned, and continuously develop expertise in Linux, GPU infrastructure, high-performance networking, storage, and automation technologies. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [10M Data Records Lost, Underwater Computing, and Psychedelic Fish - Matthias Geniar](https://www.wearedevelopers.com/videos/1908-10m-data-records-lost-underwater-computing-and-psychedelic-fish-matthias-geniar) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Eclipse Che for Infrastructure Automation](https://www.wearedevelopers.com/videos/1611-eclipse-che-for-infrastructure-automation) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Best Paying Jobs in Technology](https://www.wearedevelopers.com/magazine/256-best-paying-jobs-in-technology)