> Markdown version of [/jobs/ext/3586871-hpc-ai-workload-management-team-lead-hpc-engineer-2-3](https://www.wearedevelopers.com/jobs/ext/3586871-hpc-ai-workload-management-team-lead-hpc-engineer-2-3). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # HPC & AI Workload Management Team Lead (HPC Engineer 2/3) - **Company:** Los Alamos National Laboratory - **Location:** Los Alamos, NM, United States - **Experience:** Expert - **Salary:** $106,400.0 - $176,000.0 - **Contract:** Temporary contract - **Skills:** C (Programming Language), Artificial Intelligence, Amazon Web Services, Computing Platforms, Microsoft Azure, Bash Shell, C++ (Programming Language), Cloud Computing, Cloud Engineering, Computer Programming, Databases, Continuous Delivery, Continuous Integration, Data Visualization, Software Debugging, Linux, DevOps, Job Scheduling, Python (Programming Language), Linux System Administration, Machine Learning, Message Passing Interface, Octopus Deploy, Ansible, Software Requirements Analysis, Scripting, Google Cloud, High Performance Computing, Grafana, Infrastructure as Code (IaC), Containerization, Gitlab-ci, Kubernetes, Deployment Automation, Modeling and Simulation, Slurm, Machine Learning Operations, Terraform, Splunk - **Published:** October 5, 2026 - **Apply:** https://dejobs.org/x/x/CBAD46B34EFF4BC2A294B547A46709F0/job/ ## About the Role * Linux Administration and DevOps: Demonstrated experience administering Linux systems and applying DevOps practices to support reliable system operations. * Automation and Programming: Experience automating technical workflows and programming or scripting in languages such as Python, Bash, C, or C++. * Computing Environments: Working knowledge of one or more of the following: large-scale system administration; job scheduling and modeling and simulation workflows; or AI/ML infrastructure, deployment, and tooling. * Technical Direction and Collaboration: Ability to provide technical direction, coordinate work, and collaborate with team members, vendors, and internal partners to resolve technical issues and support production operations., * Production HPC Operations: Demonstrated experience managing production HPC environments and resolving complex operational issues. * Technical Leadership: Demonstrated accomplishments serving as a technical lead for a project, including guiding technical work and coordinating contributions from others. * Scheduling and Resource Management: Demonstrated experience administering an HPC or AI job scheduler or resource manager, such as Slurm, Moab, LSF, PBS/Torque, Grid Engine, Flux, or Run. * Parallel Runtime Systems: Experience working with the internals of Message Passing Interface (MPI) implementations, such as Open MPI or MPICH; PMIx; or a similar parallel runtime model. * Automation and Orchestration: Knowledge of automation, orchestration, and deployment tools such as Ansible, Argo CD, Kubernetes, Terraform, or similar technologies. Education/Experience for the lower level: Positions requires a Bachelor' degree in a STEM field from an accredited college and university and 4 years of related experience, typically with post-doctoral research experience at a university or national lab or equivalent experience directly related to the occupation Education/Experience for the higher level: Position requires a Master's degree in a STEM field from an accredited college or university and 6 years of relevant experience or an equivalent combination of education and experience directly related to the occupation. Desired Qualifications: * Knowledge of virtualization, containerization, and resource orchestration * Expertise in managing large, complex computing environments * Experience working with, debugging, and adding features to large code bases in languages like C and C++. * Experience with monitoring, dashboards, and data visualization (i.e. Splunk, Grafana) * Experience with OCHAMI or similar HPC cluster management frameworks * Understanding of cloud platforms and cloud-native DevOps practices (AWS, Azure, GCP) * Experience with GitLab CI, ArgoCD, or similar continuous delivery tools Work Location: The work location for this position is hybrid and is located in Los Alamos, NM. Hybrid is defined as working partially onsite/partially offsite but within 2 hours ground commute of this location. All work locations are at the discretion of management and can change at any time with appropriate notice., *Eligibility requirements: To obtain a clearance, an individual must be at least 18 years of age; U.S. citizenship is required except in very limited circumstances. See DOE Order 472.2 (https://www.directives.doe.gov/directives-documents/400-series/0472.2-border-a-chg1-ltdchg/@@images/file) for additional information. ## Description We are a small team, and each member plays a direct role in operating our production systems. We are seeking an experienced professional in HPC scheduling and resource management to lead the team in ensuring the reliable operation of our production systems, which include: * Large-Scale Modeling & Simulation (e.g., physics, engineering, climate) infrastructure * Artificial Intelligence & Machine Learning (AI/ML) deployments * Emerging Architectures for next-generation computing needs This position will be filled at either the HPC Engineer 2 or HPC Engineer 3 level, depending on the skills of the selected candidate. Additional job responsibilities, outlined below, will be assigned if the candidate is hired at the higher level. HPC Engineer 2 ($106,400-$176,000) In this role, you will: * Configure, deploy, maintain, and troubleshoot HPC job scheduling software and associated databases on existing and new production systems. * Analyze scheduling performance and optimize resource allocation for modeling and simulation, AI/ML, and other computing workloads. * Develop and maintain workflow automation, continuous integration and continuous delivery (CI/CD) pipelines, and Infrastructure as Code (IaC) solutions for testing, provisioning, configuration, and deployment. * Collaborate with vendors and internal technical teams to resolve issues, implement updates, and integrate scheduling solutions with new computing architectures. * Provide technical direction, coordinate project activities and task execution, and provide input to group management during annual performance reviews. * Participate in an on-call rotation to support production operations. HPC Engineer 3 ($128,000-$215,900) In addition to the duties outlined above, you will: * Lead the deployment and integration of new HPC systems and architectures, including automated provisioning and configuration. * Lead efforts to optimize scheduling and resource utilization across large-scale simulation and AI/ML workloads. * Work with stakeholders and vendors to define system requirements and guide technical solutions. * Evaluate and introduce new approaches to HPC system management, automation, and orchestration. * Partner with production system administrators to diagnose and resolve complex system issues affecting applications running on HPC resources. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shifting Stress to Progress— Understanding DevOps to do DevOps Better](https://www.wearedevelopers.com/videos/268-shifting-stress-to-progress-understanding-devops-to-do-devops-better) - [LLMOps-driven fine-tuning, evaluation, and inference with NVIDIA NIM & NeMo Microservices](https://www.wearedevelopers.com/videos/1582-llmops-driven-fine-tuning-evaluation-and-inference-with-nvidia-nim-nemo-microservices) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)