Linux Administrator, AI and HPC Infrastructure

Job Cloud Inc.
United States
1 day ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence User Authentication Bash Shell BIOS Ubuntu (Operating System) Cloud Computing Computer Clusters Data Centers Dynamic Host Configuration Protocol Linux RAID File Systems
+13 more
Domain Name System (DNS) Firmware Python (Programming Language) Linux System Administration Network File Systems Red Hat Enterprise Linux Ansible Shell Script High Performance Computing Data Center Networking Kubernetes Slurm Nvme

Job description

WWT is seeking a Sr. Linux Administrator, AI and HPC Infrastructure to provide senior Linux administration services across AI and HPC environments in support of a WWT customer. The environment spans GPU clusters, high-performance storage, and data center network-connected compute infrastructure. This role is intended for a hands-on operator who can stabilize production systems, resolve complex node-level failures, and improve fleet reliability at scale., * Administer large-scale Linux environments supporting AI training, inference, and HPC workloads.

  • Own deep troubleshooting of OS, kernel, boot, package, firmware, driver, filesystem, service, and resource-consumption issues across bare-metal server fleets.
  • Diagnose failures spanning BIOS, BMC, PXE, DHCP, DNS, NFS, local disk, RAID, NVMe, systemd, and GPU driver stacks.
  • Build and maintain golden images, provisioning pipelines, configuration baselines, and post-deployment validation procedures.
  • Partner with network, platform, storage, and validation teams to isolate cross-domain failures affecting cluster readiness or job execution.
  • Investigate performance anomalies involving CPU, memory, NUMA, I/O, interrupts, process scheduling, and kernel tuning.
  • Automate repeatable administration and remediation tasks with Bash and Python.
  • Produce clear runbooks, failure signatures, and escalation criteria for recurring operational issues.

Requirements

  • 7+ years delivering Linux administration in data center, cloud, AI, or HPC environments.
  • Deep expertise with RHEL, Ubuntu, Rocky, or similar enterprise Linux distributions.
  • Strong troubleshooting skill across boot flow, system logs, networking stack, authentication, service lifecycle, and hardware-software interaction.
  • Experience with GPU servers, out-of-band management, firmware coordination, and cluster node bring-up.
  • Hands-on knowledge of Ansible, PXE/iPXE, Kickstart, cloud-init, image lifecycle management, and configuration enforcement.
  • Strong shell scripting and Python-based automation capability.
  • Working knowledge of storage and network dependencies affecting Linux host health.
  • Ability to operate independently in ambiguous, high-severity production situations., * Exposure to Slurm, Kubernetes, container runtimes, or AI cluster schedulers.
  • Familiarity with GPU telemetry and health monitoring tooling (e.g., DCGM) and high-performance fabric technologies, including telemetry-driven health analysis.
  • Experience supporting validation labs or pre-production cluster certification.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

41 sec

Massive client data loss and bio-digital storage

Chris Heilmann +1 · LIVE

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

1:44 min

Career transition into cloud native and data management

Michael Cade · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all