HPC & AI Workload Management Team Lead (HPC Engineer 2/3)

Los Alamos National Laboratory
Los Alamos, NM, United States
1 day ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$106,400.0 - $176,000.0
Working hours
Regular working hours
Job source

Tech stack

C (Programming Language) Artificial Intelligence Amazon Web Services Computing Platforms Microsoft Azure Bash Shell C++ (Programming Language) Cloud Computing Cloud Engineering Computer Programming Databases Continuous Delivery
+27 more
Continuous Integration Data Visualization Software Debugging Linux DevOps Job Scheduling Python (Programming Language) Linux System Administration Machine Learning Message Passing Interface Octopus Deploy Ansible Software Requirements Analysis Scripting Google Cloud High Performance Computing Grafana Infrastructure as Code (IaC) Containerization Gitlab-ci Kubernetes Deployment Automation Modeling and Simulation Slurm Machine Learning Operations Terraform Splunk

Job description

We are a small team, and each member plays a direct role in operating our production systems. We are seeking an experienced professional in HPC scheduling and resource management to lead the team in ensuring the reliable operation of our production systems, which include:

  • Large-Scale Modeling & Simulation (e.g., physics, engineering, climate) infrastructure
  • Artificial Intelligence & Machine Learning (AI/ML) deployments
  • Emerging Architectures for next-generation computing needs

This position will be filled at either the HPC Engineer 2 or HPC Engineer 3 level, depending on the skills of the selected candidate. Additional job responsibilities, outlined below, will be assigned if the candidate is hired at the higher level.

HPC Engineer 2 ($106,400-$176,000)

In this role, you will:

  • Configure, deploy, maintain, and troubleshoot HPC job scheduling software and associated databases on existing and new production systems.
  • Analyze scheduling performance and optimize resource allocation for modeling and simulation, AI/ML, and other computing workloads.
  • Develop and maintain workflow automation, continuous integration and continuous delivery (CI/CD) pipelines, and Infrastructure as Code (IaC) solutions for testing, provisioning, configuration, and deployment.
  • Collaborate with vendors and internal technical teams to resolve issues, implement updates, and integrate scheduling solutions with new computing architectures.
  • Provide technical direction, coordinate project activities and task execution, and provide input to group management during annual performance reviews.
  • Participate in an on-call rotation to support production operations.

HPC Engineer 3 ($128,000-$215,900)

In addition to the duties outlined above, you will:

  • Lead the deployment and integration of new HPC systems and architectures, including automated provisioning and configuration.
  • Lead efforts to optimize scheduling and resource utilization across large-scale simulation and AI/ML workloads.
  • Work with stakeholders and vendors to define system requirements and guide technical solutions.
  • Evaluate and introduce new approaches to HPC system management, automation, and orchestration.
  • Partner with production system administrators to diagnose and resolve complex system issues affecting applications running on HPC resources.

Requirements

  • Linux Administration and DevOps: Demonstrated experience administering Linux systems and applying DevOps practices to support reliable system operations.
  • Automation and Programming: Experience automating technical workflows and programming or scripting in languages such as Python, Bash, C, or C++.
  • Computing Environments: Working knowledge of one or more of the following: large-scale system administration; job scheduling and modeling and simulation workflows; or AI/ML infrastructure, deployment, and tooling.
  • Technical Direction and Collaboration: Ability to provide technical direction, coordinate work, and collaborate with team members, vendors, and internal partners to resolve technical issues and support production operations., * Production HPC Operations: Demonstrated experience managing production HPC environments and resolving complex operational issues.
  • Technical Leadership: Demonstrated accomplishments serving as a technical lead for a project, including guiding technical work and coordinating contributions from others.
  • Scheduling and Resource Management: Demonstrated experience administering an HPC or AI job scheduler or resource manager, such as Slurm, Moab, LSF, PBS/Torque, Grid Engine, Flux, or Run.
  • Parallel Runtime Systems: Experience working with the internals of Message Passing Interface (MPI) implementations, such as Open MPI or MPICH; PMIx; or a similar parallel runtime model.
  • Automation and Orchestration: Knowledge of automation, orchestration, and deployment tools such as Ansible, Argo CD, Kubernetes, Terraform, or similar technologies.

Education/Experience for the lower level: Positions requires a Bachelor’ degree in a STEM field from an accredited college and university and 4 years of related experience, typically with post-doctoral research experience at a university or national lab or equivalent experience directly related to the occupation

Education/Experience for the higher level: Position requires a Master’s degree in a STEM field from an accredited college or university and 6 years of relevant experience or an equivalent combination of education and experience directly related to the occupation.

Desired Qualifications:

  • Knowledge of virtualization, containerization, and resource orchestration
  • Expertise in managing large, complex computing environments
  • Experience working with, debugging, and adding features to large code bases in languages like C and C++.
  • Experience with monitoring, dashboards, and data visualization (i.e. Splunk, Grafana)
  • Experience with OCHAMI or similar HPC cluster management frameworks
  • Understanding of cloud platforms and cloud-native DevOps practices (AWS, Azure, GCP)
  • Experience with GitLab CI, ArgoCD, or similar continuous delivery tools

Work Location: The work location for this position is hybrid and is located in Los Alamos, NM. Hybrid is defined as working partially onsite/partially offsite but within 2 hours ground commute of this location. All work locations are at the discretion of management and can change at any time with appropriate notice., *Eligibility requirements: To obtain a clearance, an individual must be at least 18 years of age; U.S. citizenship is required except in very limited circumstances. See DOE Order 472.2 (https://www.directives.doe.gov/directives-documents/400-series/0472.2-border-a-chg1-ltdchg/@@images/file) for additional information.

Benefits & conditions

Due to federal restrictions contained in the current National Defense Authorization Act, citizens of the People’s Republic of China-including the special administrative regions of Hong Kong and Macau-as well as citizens of the Islamic Republic of Iran, the Democratic People’s Republic of Korea (North Korea), and the Russian Federation, who are not Lawful Permanent Residents (“green card” holders) are prohibited from accessing facilities that support the mission, functions, and operations of national security laboratories and nuclear weapons production facilities, which includes Los Alamos National Laboratory.

Where You Will Work

Located in beautiful northern New Mexico, Los Alamos National Laboratory (LANL) is a multidisciplinary research institution engaged in strategic science on behalf of national security. Our generous benefits package includes:

§ PPO or High Deductible medical insurance with the same large nationwide network

§ Dental and vision insurance

§ Free basic life and disability insurance

§ Paid childbirth and parental leave

§ Award-winning 401(k) (6% matching plus 3.5% annually)

§ Learning opportunities and tuition assistance

§ Flexible schedules and time off (PTO and holidays)

§ Onsite gyms and wellness programs

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

2:52 min

Addressing junior hiring bottlenecks and mitigating widespread employee burnout

Hung Lee Hung Lee +3 · World Congress 2026 Europe

1:22 min

Analyzing differences between mobile and traditional backend DevOps

Mete Baydar Mete Baydar · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:27 min

Defining DevOps through its historical origins and foundational texts

Sonal Patil · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all