HPC AI Systems Administrator
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+14 more
Job description
We are seeking a high-caliber HPC AI Systems Administrator to serve as the foundational architect for our growing AI infrastructure. This role will be responsible for building a secure, scalable, and highly optimized environment to support our corporate data initiatives.
Operating at the critical intersection of infrastructure engineering and software application, you will design and maintain a robust compute platform. Your primary mission is to enable our Development Team to fine-tune and deploy production-level machine learning models smoothly, while ensuring the platform complies with enterprise-level security, governance, and data privacy policies., * Infrastructure Architecture & Management: Lead the deployment, bare-metal configuration, maintenance, and optimization of our on-premises HPC cluster and multi-GPU architecture.
- Platform Enablement: Manage the end-to-end AI software stack, including Linux OS environments, specialized GPU drivers, runtime libraries (CUDA, NCCL), and containerization platforms.
- Developer Sandbox Orchestration: Implement and maintain workload scheduling and orchestration systems (e.g., Kubernetes, Slurm, or equivalent enterprise platforms) to manage cluster resource allocation and job prioritization for engineering teams.
- Monitoring & Performance Tuning: Establish automated telemetry and monitoring dashboards to track hardware utilization, thermal limits, and memory bandwidth, ensuring maximum compute efficiency.
- Security & Governance Compliance: Operationalize strict data-at-rest and data-in-transit security baselines, ensuring the compute environment aligns with corporate Zero Trust network architecture and compliance mandates.
- Vendor Relations & Support: Act as the primary technical interface for high-end hardware vendors and system integrators to manage system updates and platform maintenance.
Requirements
- Experience: 3+ years of dedicated systems administration experience managing Linux-based High-Performance Computing (HPC) environments or enterprise-scale GPU infrastructure.
- Technical Fluency: Hands-on experience configuring and maintaining modern enterprise GPU hardware (such as NVIDIA Ampere or Hopper architecture) in a data center context.
- Software & Tooling Mastery: Deep expertise in Linux system engineering, container technologies (Docker, Apptainer/Singularity), and cluster resource management.
- Networking & Storage: Solid baseline knowledge of high-throughput networking fabrics (e.g., InfiniBand/RoCE) and parallel or distributed enterprise storage systems.
- Education: Bachelor’s degree in computer science, Computer Engineering, System Administration, or equivalent practical industry experience.
Preferred Attributes
- Relevant professional certifications in enterprise AI infrastructure, virtualization, or cloud/hybrid architecture solutions.
- Familiarity with the infrastructure requirements supporting modern AI frameworks, machine learning lifecycles, or Large Language Model (LLM) fine-tuning pipelines.
Benefits & conditions
-
Direct access to cutting-edge, top-tier enterprise compute infrastructure. A collaborative environment with highly defined responsibilities and strategic, top-down execution support.
-
Competitive salary, comprehensive benefits package, and professional development support.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
MLOps And AI Driven Development
7 Cloud Computing Trends Coming in 2025 for Developers
MLOps – What’s the deal behind it?