> Markdown version of [/jobs/ext/1880889-ai-hpc-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/1880889-ai-hpc-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI HPC Infrastructure Engineer - **Company:** Analysis Group - **Location:** Boston, MA, United States - **Experience:** Expert - **Salary:** $150,000.0 - $170,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Nvidia CUDA, Linux, Distributed Computing Environment, Emulators, General Parallel File Systems, Job Scheduling, Python (Programming Language), RStudio, Remote Access Technology, Ansible, Tensorflow, Pytorch, Large Language Models, Containerization, Kubernetes, Information Technology, Slurm, Machine Learning Operations, Hardware Infrastructure, Docker - **Published:** August 1, 2026 - **Apply:** https://jobs.military.com/career/295752/ai-hpc-infrastructure-engineer-massachusetts-ma-boston ## About the Role * Bachelor's degree required; degree in computer science, electrical engineering, or a related field preferred.\n * A minimum of 5 years of experience as a hands-on Linux Systems Administrator in a research, HPC, or production setting.\n * An ideal candidate will have 5 to 10 years of substantive relevant experience. \n * Experience managing Posit Workbench (RStudio Server Pro), Python, and R environments; strong Posit Workbench administration experience is a significant plus.\n * Experience with SLURM, Platform LSF, or other job schedulers required; experience scheduling GPU resources strongly preferred.\n * Hands-on experience with NVIDIA GPU infrastructure and software stack (CUDA, cuDNN, NCCL, NVIDIA GPU Operator) strongly preferred.\n * Experience with Kubernetes and container orchestration for AI/ML workloads highly desired.\n * Familiarity with ML/AI frameworks (PyTorch, TensorFlow) and distributed training patterns highly desired.\n * Experience with MLOps tooling (MLflow, Kubeflow, Weights & Biases, or similar) is a plus.\n * Experience with Bright Cluster Manager is highly desired.\n * Experience with Ansible is highly desired.\n * Experience with containerization (Docker, Singularity/Apptainer) is highly desired.\n * Proficiency with remote access technologies and tools such as RDP, SSH, and emulation software\n * Hands-on experience with GPFS (IBM Spectrum Scale) required.\n * Demonstrated experience tuning LLM training and/or inference performance (e.g., batching, quantization, KV-cache management, parallelism strategies) required.\n * Experience with AI Gateways (e.g., LiteLLM, Kong AI Gateway, Portkey, or similar) is a very nice to have.\n * Excellent hardware troubleshooting experience, including GPU-specific diagnostics.\n * Knowledge of applicable data privacy practices and laws.\n * Strong interpersonal, written, and oral communication skills.\n * Highly self-motivated and directed, with keen attention to detail.\n * Proven analytical and problem-solving abilities.\n * Strong customer service orientation.\n * Experience working in a collaborative environment.\n * An inclusive and growth-oriented mindset, strong interpersonal skills, and an ability to work across functions.\n * To the extent permitted by applicable law, eligible candidates must be authorized to work in the United States, without sponsorship or restriction, now and in the future.\n ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)