> Markdown version of [/jobs/ext/2499546-linux-administrator-ai-hpc-infrastructure](https://www.wearedevelopers.com/jobs/ext/2499546-linux-administrator-ai-hpc-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Linux Administrator- AI,HPC Infrastructure - **Company:** Job Cloud Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Temporary contract - **Skills:** Artificial Intelligence, User Authentication, Bash Shell, BIOS, Ubuntu (Operating System), Cloud Computing, Data Centers, Dynamic Host Configuration Protocol, Linux, RAID, File Systems, Domain Name System (DNS), Firmware, Python (Programming Language), Linux System Administration, Network File Systems, Red Hat Enterprise Linux, Ansible, Shell Script, High Performance Computing, Kubernetes, Slurm, Nvme - **Published:** August 15, 2026 - **Apply:** https://www.dice.com/job-detail/30e5c805-9dce-4f50-98ef-d821b250dfc0 ## About the Role * 7+ years delivering Linux administration in data center, cloud, AI, or HPC environments. * Deep expertise with RHEL, Ubuntu, Rocky, or similar enterprise Linux distributions. * Strong troubleshooting skill across boot flow, system logs, networking stack, authentication, service lifecycle, and hardware-software interaction. * Experience with GPU servers, out-of-band management, firmware coordination, and cluster node bring-up. * Hands-on knowledge of Ansible, PXE/iPXE, Kickstart, cloud-init, image lifecycle management, and configuration enforcement. * Strong shell scripting and Python-based automation capability. * Working knowledge of storage and network dependencies affecting Linux host health. * Ability to operate independently in ambiguous, high-severity production situations., * Exposure to Slurm, Kubernetes, container runtimes, or AI cluster schedulers. * Familiarity with GPU telemetry and health monitoring tooling (e.g., DCGM) and high-performance fabric technologies, including telemetry-driven health analysis. * Experience supporting validation labs or pre-production cluster certification. ## Description * Administer large-scale Linux environments supporting AI training, inference, and HPC workloads. * Own deep troubleshooting of OS, kernel, boot, package, firmware, driver, filesystem, service, and resource-consumption issues across bare-metal server fleets. * Diagnose failures spanning BIOS, BMC, PXE, DHCP, DNS, NFS, local disk, RAID, NVMe, systemd, and GPU driver stacks. * Build and maintain golden images, provisioning pipelines, configuration baselines, and post-deployment validation procedures. * Partner with network, platform, storage, and validation teams to isolate cross-domain failures affecting cluster readiness or job execution. * Investigate performance anomalies involving CPU, memory, NUMA, I/O, interrupts, process scheduling, and kernel tuning. * Automate repeatable administration and remediation tasks with Bash and Python. * Produce clear runbooks, failure signatures, and escalation criteria for recurring operational issues. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [10M Data Records Lost, Underwater Computing, and Psychedelic Fish - Matthias Geniar](https://www.wearedevelopers.com/videos/1908-10m-data-records-lost-underwater-computing-and-psychedelic-fish-matthias-geniar) - [How to Submit CFPs and Get into Public Speaking - Moran Weber](https://www.wearedevelopers.com/videos/2112-how-to-submit-cfps-and-get-into-public-speaking-moran-weber) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)