HPC AI Systems Administrator

MRE CONSULTING, LTD.
Houston, TX, United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Software Applications Computing Platforms Systems Engineering Nvidia CUDA Computer Engineering Data Centers Linux InfiniBand Machine Learning Performance Tuning Zero Trust Network Access
+14 more
Virtualization Technology AI Infrastructure Enterprise Data Management Graphics Processing Unit (GPU) Large Language Models IT Architecture Containerization Kubernetes Information Technology Bare Metal Data Management Slurm Hardware Infrastructure Docker

Job description

We are seeking a high-caliber HPC AI Systems Administrator to serve as the foundational architect for our growing AI infrastructure. This role will be responsible for building a secure, scalable, and highly optimized environment to support our corporate data initiatives.

Operating at the critical intersection of infrastructure engineering and software application, you will design and maintain a robust compute platform. Your primary mission is to enable our Development Team to fine-tune and deploy production-level machine learning models smoothly, while ensuring the platform complies with enterprise-level security, governance, and data privacy policies., * Infrastructure Architecture & Management: Lead the deployment, bare-metal configuration, maintenance, and optimization of our on-premises HPC cluster and multi-GPU architecture.

  • Platform Enablement: Manage the end-to-end AI software stack, including Linux OS environments, specialized GPU drivers, runtime libraries (CUDA, NCCL), and containerization platforms.
  • Developer Sandbox Orchestration: Implement and maintain workload scheduling and orchestration systems (e.g., Kubernetes, Slurm, or equivalent enterprise platforms) to manage cluster resource allocation and job prioritization for engineering teams.
  • Monitoring & Performance Tuning: Establish automated telemetry and monitoring dashboards to track hardware utilization, thermal limits, and memory bandwidth, ensuring maximum compute efficiency.
  • Security & Governance Compliance: Operationalize strict data-at-rest and data-in-transit security baselines, ensuring the compute environment aligns with corporate Zero Trust network architecture and compliance mandates.
  • Vendor Relations & Support: Act as the primary technical interface for high-end hardware vendors and system integrators to manage system updates and platform maintenance.

Requirements

  • Experience: 3+ years of dedicated systems administration experience managing Linux-based High-Performance Computing (HPC) environments or enterprise-scale GPU infrastructure.
  • Technical Fluency: Hands-on experience configuring and maintaining modern enterprise GPU hardware (such as NVIDIA Ampere or Hopper architecture) in a data center context.
  • Software & Tooling Mastery: Deep expertise in Linux system engineering, container technologies (Docker, Apptainer/Singularity), and cluster resource management.
  • Networking & Storage: Solid baseline knowledge of high-throughput networking fabrics (e.g., InfiniBand/RoCE) and parallel or distributed enterprise storage systems.
  • Education: Bachelor’s degree in computer science, Computer Engineering, System Administration, or equivalent practical industry experience.

Preferred Attributes

  • Relevant professional certifications in enterprise AI infrastructure, virtualization, or cloud/hybrid architecture solutions.
  • Familiarity with the infrastructure requirements supporting modern AI frameworks, machine learning lifecycles, or Large Language Model (LLM) fine-tuning pipelines.

Benefits & conditions

  • Direct access to cutting-edge, top-tier enterprise compute infrastructure. A collaborative environment with highly defined responsibilities and strategic, top-down execution support.

  • Competitive salary, comprehensive benefits package, and professional development support.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:58 min

Verifying hardware access and exploring AI inference scaling

Piotr Zaniewski Piotr Zaniewski · World Congress 2026 Europe

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all