High Performance Computing Engineer for Workflow Modernization (HPC Engineer 2/3)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+21 more
Job description
The HPC Division is seeking an engineer to modernize how our Workload Management and Consulting teams do their own work, using AI to make our tools, documentation, and support processes more effective. This is a role between infrastructure and end users - you’ll build capabilities that make both our team and our users more productive.
This position will be filled at either the HPC Engineer 2 or HPC Engineer 3 level, depending on the skills of the selected candidate.
Core responsibilities:
- Identify which of the group’s existing processes - documentation generation and upkeep, troubleshooting guides - are good candidates for AI-assisted automation, and build that automation. This includes surfacing relevant system and knowledge base content to assist consultants investigating complex user issues.
- Enable AI workloads on LANL’s HPC systems through investigation and deployment of relevant technologies, such as:
- Slinky, a next-generation collection of projects bridging HPC and AI workloads.
- RunAI, a scheduler for GPU-centric AI/ML workloads, including capacity sharing and integration with existing HPC resources.
- Develop and deploy tools on OpenShift, for both direct user consumption and internal division use, including infrastructure that gives users visibility into their workloads.
- Provide technical support and documentation for the tools and workflows you build, and collaborate with vendors and open-source communities on relevant scheduling, orchestration, and AI tooling.
The tools you build will serve two audiences: HPC users trying to get work done, and HPC staff trying to support them.
HPC Engineer 2 (106,400-176,000/year)
At this level, responsibilities include:
- Building and iterating on AI-assisted tools for documentation and support workflows, under guidance from senior staff.
- Assisting with evaluation, deployment, and support of Slinky and RunAI in test and production environments.
- Developing and deploying tools on OpenShift.
- Providing day-to-day technical support and troubleshooting for users of these systems.
- Contributing to documentation and knowledge base content, including using the tools you help build.
- Collaborating with the Workload Management and Consulting teams to identify workflow pain points worth solving.
This level applies standard principles and techniques, builds familiarity with new AI tooling and scheduling/orchestration technologies, and exercises independent judgment on well-scoped problems.
HPC Engineer 3 (128,000-215,900/year)
In addition to the above, this level includes:
- Technical ownership of the group’s AI-assisted workflow modernization efforts - deciding what to build, how to build it, and how to evaluate whether it’s actually helping.
- Designing more complex tooling that combines AI-assisted automation with observability into user and staff workflows.
- Technical ownership of significant pieces of the Slinky and/or RunAI evaluation and deployment effort, drawing on hands-on experience with HPC workload management to guide design and integration decisions.
- Root-cause troubleshooting of complex issues spanning scheduling, orchestration, and the internal tools you build.
- Representing this work in technical discussions with vendors, other HPC teams, and the broader HPC community.
- Mentoring junior staff and students on AI tooling, scheduling systems, and container orchestration.
This level requires deep technical expertise in applying AI tooling to real operational problems, alongside solid grounding in scheduling and container orchestration.
About the Organization
LANL’s High Performance Computing Division supports the Laboratory’s mission by managing a world-class supercomputing center, supporting stockpile stewardship for NNSA/DOE and accelerating scientific discovery through some of the world’s largest supercomputers.
Within HPC, the Environments Group (HPC-ENV) is responsible for the user experience of our supercomputers. This position sits between 2 teams within HPC-ENV:
- Workload Management is responsible for the schedulers, resource managers, and orchestration layers that determine how and when work runs on LANL’s supercomputers, including evaluation of new technologies.
- Consulting is the single point of contact for HPC customers, providing direct technical support and coordinating institutional knowledge across the division., * LLM tooling and frameworks (LangChain, LlamaIndex, vector databases, local/hosted model serving, etc.).
- Slurm or other HPC workload managers.
- RunAI or comparable Kubernetes-native AI/ML scheduling platforms.
- OpenShift or Kubernetes application development and deployment.
- Documentation tooling (static site generators, wikis, knowledge base platforms) and how to automate or augment them.
- Configuration management tools (Ansible, Chef, etc.).
- CI/CD pipelines (GitLab CI, GitHub Actions, etc.).
- Monitoring/observability stacks (Prometheus, Grafana, ELK, etc.).
Institutional Growth
- Skills or perspectives relevant to the role’s goals not listed above, technical or otherwise.
- Interest in making complex systems and processes usable, not just functional, for the people who depend on them.
Requirements
- Effective written and oral communication skills.
- Experience with UNIX/Linux systems administration and command-line tools.
- Experience applying AI/ML tools to practical problems - document generation/summarization, classification, retrieval-augmented systems, or similar.
- Experience with containers (Docker, Podman, Charliecloud, Singularity/Apptaine, etc..) and container orchestration (OpenShift, Rancher, etc..).
- Scripting ability (Python, Bash, or similar) for automation and tool development.
- Ability to communicate across both infrastructure and end-user contexts.
Additional Requirements for HPC Engineer 3:
- Demonstrated depth of experience designing and deploying AI-assisted tooling for real operational or support workflows.
- Experience evaluating new technologies and producing concrete deployment recommendations.
- Hands-on experience with HPC systems and workload management.
Education/Experience:
HPC Engineer 2: Bachelor’s degree in Computer Science, Computer Engineering, or related field, and 3 years of relevant experience in HPC, scalable AI computing, or data center environments, or equivalent combination of education and experience.
HPC Engineer 3: Bachelor’s degree in Computer Science, Computer Engineering, or related field, and 6 years of relevant experience in HPC, scalable AI computing, or data center environments, or equivalent combination of education and experience., * Experience building tools around large language models - retrieval-augmented generation, fine-tuning, prompt engineering, agentic workflows, or similar - applied to documentation or support use cases.
- Experience with HPC or cluster workload managers/schedulers (Slurm, PBS, LSF, or similar).
- GPU scheduling and resource sharing in HPC or cloud environments.
- Building observability tooling (metrics, logging, dashboards) for distributed or scientific computing workflows.
- Customer service or public-facing technical support experience.
- Evaluating software/systems against performance or usability metrics and producing improvement recommendations., *Eligibility requirements: To obtain a clearance, an individual must be at least 18 years of age; U.S. citizenship is required except in very limited circumstances. See DOE Order 472.2 (https://www.directives.doe.gov/directives-documents/400-series/0472.2-border-a-chg1-ltdchg/@@images/file) for additional information.
Benefits & conditions
Due to federal restrictions contained in the current National Defense Authorization Act, citizens of the People’s Republic of China-including the special administrative regions of Hong Kong and Macau-as well as citizens of the Islamic Republic of Iran, the Democratic People’s Republic of Korea (North Korea), and the Russian Federation, who are not Lawful Permanent Residents (“green card” holders) are prohibited from accessing facilities that support the mission, functions, and operations of national security laboratories and nuclear weapons production facilities, which includes Los Alamos National Laboratory.
Where You Will Work
Located in beautiful northern New Mexico, Los Alamos National Laboratory (LANL) is a multidisciplinary research institution engaged in strategic science on behalf of national security. Our generous benefits package includes:
§ PPO or High Deductible medical insurance with the same large nationwide network
§ Dental and vision insurance
§ Free basic life and disability insurance
§ Paid childbirth and parental leave
§ Award-winning 401(k) (6% matching plus 3.5% annually)
§ Learning opportunities and tuition assistance
§ Flexible schedules and time off (PTO and holidays)
§ Onsite gyms and wellness programs
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Top 6 Hackathons for Developers in 2023
7 Cloud Computing Trends Coming in 2025 for Developers
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
Dev Digest 121 - AI goes offline