HPC & AI Workload Management Team Lead (HPC Engineer 2/3)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+27 more
Job description
We are a small team, and each member plays a direct role in operating our production systems. We are seeking an experienced professional in HPC scheduling and resource management to lead the team in ensuring the reliable operation of our production systems, which include:
- Large-Scale Modeling & Simulation (e.g., physics, engineering, climate) infrastructure
- Artificial Intelligence & Machine Learning (AI/ML) deployments
- Emerging Architectures for next-generation computing needs
This position will be filled at either the HPC Engineer 2 or HPC Engineer 3 level, depending on the skills of the selected candidate. Additional job responsibilities, outlined below, will be assigned if the candidate is hired at the higher level.
HPC Engineer 2 ($106,400-$176,000)
In this role, you will:
- Configure, deploy, maintain, and troubleshoot HPC job scheduling software and associated databases on existing and new production systems.
- Analyze scheduling performance and optimize resource allocation for modeling and simulation, AI/ML, and other computing workloads.
- Develop and maintain workflow automation, continuous integration and continuous delivery (CI/CD) pipelines, and Infrastructure as Code (IaC) solutions for testing, provisioning, configuration, and deployment.
- Collaborate with vendors and internal technical teams to resolve issues, implement updates, and integrate scheduling solutions with new computing architectures.
- Provide technical direction, coordinate project activities and task execution, and provide input to group management during annual performance reviews.
- Participate in an on-call rotation to support production operations.
HPC Engineer 3 ($128,000-$215,900)
In addition to the duties outlined above, you will:
- Lead the deployment and integration of new HPC systems and architectures, including automated provisioning and configuration.
- Lead efforts to optimize scheduling and resource utilization across large-scale simulation and AI/ML workloads.
- Work with stakeholders and vendors to define system requirements and guide technical solutions.
- Evaluate and introduce new approaches to HPC system management, automation, and orchestration.
- Partner with production system administrators to diagnose and resolve complex system issues affecting applications running on HPC resources.
Requirements
- Linux Administration and DevOps: Demonstrated experience administering Linux systems and applying DevOps practices to support reliable system operations.
- Automation and Programming: Experience automating technical workflows and programming or scripting in languages such as Python, Bash, C, or C++.
- Computing Environments: Working knowledge of one or more of the following: large-scale system administration; job scheduling and modeling and simulation workflows; or AI/ML infrastructure, deployment, and tooling.
- Technical Direction and Collaboration: Ability to provide technical direction, coordinate work, and collaborate with team members, vendors, and internal partners to resolve technical issues and support production operations., * Production HPC Operations: Demonstrated experience managing production HPC environments and resolving complex operational issues.
- Technical Leadership: Demonstrated accomplishments serving as a technical lead for a project, including guiding technical work and coordinating contributions from others.
- Scheduling and Resource Management: Demonstrated experience administering an HPC or AI job scheduler or resource manager, such as Slurm, Moab, LSF, PBS/Torque, Grid Engine, Flux, or Run.
- Parallel Runtime Systems: Experience working with the internals of Message Passing Interface (MPI) implementations, such as Open MPI or MPICH; PMIx; or a similar parallel runtime model.
- Automation and Orchestration: Knowledge of automation, orchestration, and deployment tools such as Ansible, Argo CD, Kubernetes, Terraform, or similar technologies.
Education/Experience for the lower level: Positions requires a Bachelor’ degree in a STEM field from an accredited college and university and 4 years of related experience, typically with post-doctoral research experience at a university or national lab or equivalent experience directly related to the occupation
Education/Experience for the higher level: Position requires a Master’s degree in a STEM field from an accredited college or university and 6 years of relevant experience or an equivalent combination of education and experience directly related to the occupation.
Desired Qualifications:
- Knowledge of virtualization, containerization, and resource orchestration
- Expertise in managing large, complex computing environments
- Experience working with, debugging, and adding features to large code bases in languages like C and C++.
- Experience with monitoring, dashboards, and data visualization (i.e. Splunk, Grafana)
- Experience with OCHAMI or similar HPC cluster management frameworks
- Understanding of cloud platforms and cloud-native DevOps practices (AWS, Azure, GCP)
- Experience with GitLab CI, ArgoCD, or similar continuous delivery tools
Work Location: The work location for this position is hybrid and is located in Los Alamos, NM. Hybrid is defined as working partially onsite/partially offsite but within 2 hours ground commute of this location. All work locations are at the discretion of management and can change at any time with appropriate notice., *Eligibility requirements: To obtain a clearance, an individual must be at least 18 years of age; U.S. citizenship is required except in very limited circumstances. See DOE Order 472.2 (https://www.directives.doe.gov/directives-documents/400-series/0472.2-border-a-chg1-ltdchg/@@images/file) for additional information.
Benefits & conditions
Due to federal restrictions contained in the current National Defense Authorization Act, citizens of the People’s Republic of China-including the special administrative regions of Hong Kong and Macau-as well as citizens of the Islamic Republic of Iran, the Democratic People’s Republic of Korea (North Korea), and the Russian Federation, who are not Lawful Permanent Residents (“green card” holders) are prohibited from accessing facilities that support the mission, functions, and operations of national security laboratories and nuclear weapons production facilities, which includes Los Alamos National Laboratory.
Where You Will Work
Located in beautiful northern New Mexico, Los Alamos National Laboratory (LANL) is a multidisciplinary research institution engaged in strategic science on behalf of national security. Our generous benefits package includes:
§ PPO or High Deductible medical insurance with the same large nationwide network
§ Dental and vision insurance
§ Free basic life and disability insurance
§ Paid childbirth and parental leave
§ Award-winning 401(k) (6% matching plus 3.5% annually)
§ Learning opportunities and tuition assistance
§ Flexible schedules and time off (PTO and holidays)
§ Onsite gyms and wellness programs
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Top 6 Hackathons for Developers in 2023
7 Cloud Computing Trends Coming in 2025 for Developers
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
How to Become an AI Engineer