> Markdown version of [/jobs/ext/3115383-high-performance-computing-engineer-for-workflow-modernization-hpc-engineer-2-3](https://www.wearedevelopers.com/jobs/ext/3115383-high-performance-computing-engineer-for-workflow-modernization-hpc-engineer-2-3). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # High Performance Computing Engineer for Workflow Modernization (HPC Engineer 2/3) - **Company:** Los Alamos National Laboratory - **Location:** Los Alamos, NM, United States - **Salary:** $106,400.0 - $176,000.0 - **Contract:** Temporary contract - **Skills:** Artificial Intelligence, Bash Shell, Command-Line Interface, Computer Engineering, Data Centers, Github, Python (Programming Language), Open Source Technology, OpenShift, Ansible, Prometheus, Scientific Computating, Software Engineering, Supercomputing, Datadog, Planning Software, Data Logging, Scripting, Cloud Platform System, High Performance Computing, Retrieval-Augmented Generation, Large Language Models, Grafana, Prompt Engineering, Gitlab-ci, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Rancher, Wikis, Slurm, Machine Learning Operations, Docker - **Published:** September 27, 2026 - **Apply:** https://dejobs.org/x/x/E842D2CBEE5C44A7A5B5345E9E34B1B8/job/ ## About the Role * Effective written and oral communication skills. * Experience with UNIX/Linux systems administration and command-line tools. * Experience applying AI/ML tools to practical problems - document generation/summarization, classification, retrieval-augmented systems, or similar. * Experience with containers (Docker, Podman, Charliecloud, Singularity/Apptaine, etc..) and container orchestration (OpenShift, Rancher, etc..). * Scripting ability (Python, Bash, or similar) for automation and tool development. * Ability to communicate across both infrastructure and end-user contexts. Additional Requirements for HPC Engineer 3: * Demonstrated depth of experience designing and deploying AI-assisted tooling for real operational or support workflows. * Experience evaluating new technologies and producing concrete deployment recommendations. * Hands-on experience with HPC systems and workload management. Education/Experience: HPC Engineer 2: Bachelor's degree in Computer Science, Computer Engineering, or related field, and 3 years of relevant experience in HPC, scalable AI computing, or data center environments, or equivalent combination of education and experience. HPC Engineer 3: Bachelor's degree in Computer Science, Computer Engineering, or related field, and 6 years of relevant experience in HPC, scalable AI computing, or data center environments, or equivalent combination of education and experience., * Experience building tools around large language models - retrieval-augmented generation, fine-tuning, prompt engineering, agentic workflows, or similar - applied to documentation or support use cases. * Experience with HPC or cluster workload managers/schedulers (Slurm, PBS, LSF, or similar). * GPU scheduling and resource sharing in HPC or cloud environments. * Building observability tooling (metrics, logging, dashboards) for distributed or scientific computing workflows. * Customer service or public-facing technical support experience. * Evaluating software/systems against performance or usability metrics and producing improvement recommendations., *Eligibility requirements: To obtain a clearance, an individual must be at least 18 years of age; U.S. citizenship is required except in very limited circumstances. See DOE Order 472.2 (https://www.directives.doe.gov/directives-documents/400-series/0472.2-border-a-chg1-ltdchg/@@images/file) for additional information. ## Description The HPC Division is seeking an engineer to modernize how our Workload Management and Consulting teams do their own work, using AI to make our tools, documentation, and support processes more effective. This is a role between infrastructure and end users - you'll build capabilities that make both our team and our users more productive. This position will be filled at either the HPC Engineer 2 or HPC Engineer 3 level, depending on the skills of the selected candidate. Core responsibilities: * Identify which of the group's existing processes - documentation generation and upkeep, troubleshooting guides - are good candidates for AI-assisted automation, and build that automation. This includes surfacing relevant system and knowledge base content to assist consultants investigating complex user issues. * Enable AI workloads on LANL's HPC systems through investigation and deployment of relevant technologies, such as: * Slinky, a next-generation collection of projects bridging HPC and AI workloads. * RunAI, a scheduler for GPU-centric AI/ML workloads, including capacity sharing and integration with existing HPC resources. * Develop and deploy tools on OpenShift, for both direct user consumption and internal division use, including infrastructure that gives users visibility into their workloads. * Provide technical support and documentation for the tools and workflows you build, and collaborate with vendors and open-source communities on relevant scheduling, orchestration, and AI tooling. The tools you build will serve two audiences: HPC users trying to get work done, and HPC staff trying to support them. HPC Engineer 2 (106,400-176,000/year) At this level, responsibilities include: * Building and iterating on AI-assisted tools for documentation and support workflows, under guidance from senior staff. * Assisting with evaluation, deployment, and support of Slinky and RunAI in test and production environments. * Developing and deploying tools on OpenShift. * Providing day-to-day technical support and troubleshooting for users of these systems. * Contributing to documentation and knowledge base content, including using the tools you help build. * Collaborating with the Workload Management and Consulting teams to identify workflow pain points worth solving. This level applies standard principles and techniques, builds familiarity with new AI tooling and scheduling/orchestration technologies, and exercises independent judgment on well-scoped problems. HPC Engineer 3 (128,000-215,900/year) In addition to the above, this level includes: * Technical ownership of the group's AI-assisted workflow modernization efforts - deciding what to build, how to build it, and how to evaluate whether it's actually helping. * Designing more complex tooling that combines AI-assisted automation with observability into user and staff workflows. * Technical ownership of significant pieces of the Slinky and/or RunAI evaluation and deployment effort, drawing on hands-on experience with HPC workload management to guide design and integration decisions. * Root-cause troubleshooting of complex issues spanning scheduling, orchestration, and the internal tools you build. * Representing this work in technical discussions with vendors, other HPC teams, and the broader HPC community. * Mentoring junior staff and students on AI tooling, scheduling systems, and container orchestration. This level requires deep technical expertise in applying AI tooling to real operational problems, alongside solid grounding in scheduling and container orchestration. About the Organization LANL's High Performance Computing Division supports the Laboratory's mission by managing a world-class supercomputing center, supporting stockpile stewardship for NNSA/DOE and accelerating scientific discovery through some of the world's largest supercomputers. Within HPC, the Environments Group (HPC-ENV) is responsible for the user experience of our supercomputers. This position sits between 2 teams within HPC-ENV: * Workload Management is responsible for the schedulers, resource managers, and orchestration layers that determine how and when work runs on LANL's supercomputers, including evaluation of new technologies. * Consulting is the single point of contact for HPC customers, providing direct technical support and coordinating institutional knowledge across the division., * LLM tooling and frameworks (LangChain, LlamaIndex, vector databases, local/hosted model serving, etc.). * Slurm or other HPC workload managers. * RunAI or comparable Kubernetes-native AI/ML scheduling platforms. * OpenShift or Kubernetes application development and deployment. * Documentation tooling (static site generators, wikis, knowledge base platforms) and how to automate or augment them. * Configuration management tools (Ansible, Chef, etc.). * CI/CD pipelines (GitLab CI, GitHub Actions, etc.). * Monitoring/observability stacks (Prometheus, Grafana, ELK, etc.). Institutional Growth * Skills or perspectives relevant to the role's goals not listed above, technical or otherwise. * Interest in making complex systems and processes usable, not just functional, for the people who depend on them. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Living Documentation That Can't Die](https://www.wearedevelopers.com/videos/2025-living-documentation-that-can-t-die) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [The Impact of AI on Game Development and the Industry](https://www.wearedevelopers.com/videos/100297-the-impact-of-ai-on-game-development-and-the-industry) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)