> Markdown version of [/jobs/ext/2383491-hpc-ai-systems-administrator](https://www.wearedevelopers.com/jobs/ext/2383491-hpc-ai-systems-administrator). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # HPC AI Systems Administrator - **Company:** MRE CONSULTING, LTD. - **Location:** Houston, TX, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Software Applications, Computing Platforms, Systems Engineering, Nvidia CUDA, Computer Engineering, Data Centers, Linux, InfiniBand, Machine Learning, Performance Tuning, Zero Trust Network Access, Virtualization Technology, AI Infrastructure, Enterprise Data Management, Graphics Processing Unit (GPU), Large Language Models, IT Architecture, Containerization, Kubernetes, Information Technology, Bare Metal, Data Management, Slurm, Hardware Infrastructure, Docker - **Published:** August 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=1e4657c13813794b ## About the Role * Experience: 3+ years of dedicated systems administration experience managing Linux-based High-Performance Computing (HPC) environments or enterprise-scale GPU infrastructure. * Technical Fluency: Hands-on experience configuring and maintaining modern enterprise GPU hardware (such as NVIDIA Ampere or Hopper architecture) in a data center context. * Software & Tooling Mastery: Deep expertise in Linux system engineering, container technologies (Docker, Apptainer/Singularity), and cluster resource management. * Networking & Storage: Solid baseline knowledge of high-throughput networking fabrics (e.g., InfiniBand/RoCE) and parallel or distributed enterprise storage systems. * Education: Bachelor's degree in computer science, Computer Engineering, System Administration, or equivalent practical industry experience. Preferred Attributes * Relevant professional certifications in enterprise AI infrastructure, virtualization, or cloud/hybrid architecture solutions. * Familiarity with the infrastructure requirements supporting modern AI frameworks, machine learning lifecycles, or Large Language Model (LLM) fine-tuning pipelines. ## Description We are seeking a high-caliber HPC AI Systems Administrator to serve as the foundational architect for our growing AI infrastructure. This role will be responsible for building a secure, scalable, and highly optimized environment to support our corporate data initiatives. Operating at the critical intersection of infrastructure engineering and software application, you will design and maintain a robust compute platform. Your primary mission is to enable our Development Team to fine-tune and deploy production-level machine learning models smoothly, while ensuring the platform complies with enterprise-level security, governance, and data privacy policies., * Infrastructure Architecture & Management: Lead the deployment, bare-metal configuration, maintenance, and optimization of our on-premises HPC cluster and multi-GPU architecture. * Platform Enablement: Manage the end-to-end AI software stack, including Linux OS environments, specialized GPU drivers, runtime libraries (CUDA, NCCL), and containerization platforms. * Developer Sandbox Orchestration: Implement and maintain workload scheduling and orchestration systems (e.g., Kubernetes, Slurm, or equivalent enterprise platforms) to manage cluster resource allocation and job prioritization for engineering teams. * Monitoring & Performance Tuning: Establish automated telemetry and monitoring dashboards to track hardware utilization, thermal limits, and memory bandwidth, ensuring maximum compute efficiency. * Security & Governance Compliance: Operationalize strict data-at-rest and data-in-transit security baselines, ensuring the compute environment aligns with corporate Zero Trust network architecture and compliance mandates. * Vendor Relations & Support: Act as the primary technical interface for high-end hardware vendors and system integrators to manage system updates and platform maintenance. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)