HPC Engineer 2/3
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+28 more
Job description
Join the High Performance Computing Environments Group (HPC-ENV) at Los Alamos National Laboratory, where we manage and operate advanced large-scale computing infrastructure supporting a diverse range of critical workloads. The HPC Consulting team provides direct support to HPC customers to enhance the productivity of the user community by providing quality technical support for a wide range of workloads, including:
- Large-Scale Modeling & Simulation (e.g., physics, engineering, climate)
- Artificial Intelligence & Machine Learning (AI/ML) deployments
- Emerging software for next-generation computing needs
Team Responsibilities
Our Consultants are a user-facing team who provide support and engage with LANL customers via documentation, consulting, and training. As HPC Consultants, we are a single point of contact for HPC systems and deliver customer-focused services, tools, and software. The Division’s goal is to create an effective HPC environment in which scientists can be as productive as possible on world-class supercomputing systems.
The HPC Division supports the Los Alamos National Laboratory (LANL) mission by managing a world-class supercomputing center. We support stockpile stewardship for NNSA/DOE and accelerate scientific discovery for scientists. We integrate and support some of the world’s largest supercomputers during an exciting time in computing, with a focus on traditional large-scale simulations, data science, artificial intelligence, and machine learning.
HPC Engineer 2 ($106,400-176,000)
- Direct Customer Support: Provide expert guidance and consultation to scientists, engineers, and other HPC users to enhance their workflows and enable effective use of HPC systems through phone, email, ticketing systems, and in-person consultations.
- Documentation and Training: Develop user guides, best practices, and training materials to help the user community get the most from the HPC environment.
- System Efficiency: Work with HPC system administrators to identify issues and improve overall system performance and resource utilization.
- Independent Problem Solving: Work independently to implement customer solutions to current problems and enhance the HPC user environment.
- External Representation: Represent LANL HPC Consulting at workshops, meetings, and conferences with other HPC and NNSA/DOE sites.
- HPC Application Support: Provide in-depth support for users of HPC systems, including installation, configuration, and troubleshooting of scientific and engineering applications.
- Performance Optimization: Analyze and optimize the performance of parallel applications and workflows to ensure efficient use of HPC resources.
- Code Profiling and Tuning: Collaborate with software developers and users to profile and tune applications for performance, memory usage, and scalability on HPC platforms.
HPC Engineer 3 ($128,000-215,900)
In addition to the duties outlined above, a successful HPC Engineer 3 candidate will be required to:
- Advanced Application Support: Serve as the go-to technical resource for the most complex scientific and engineering application issues, including deep debugging of parallel codes, build failures, and cross-platform portability problems.
- Performance Engineering: Conduct advanced profiling, tuning, and optimization of parallel applications across CPU and GPU architectures, identifying bottlenecks in memory, I/O, communication, and scalability.
- Software Stack Expertise: Maintain deep working knowledge of compilers (NVIDIA HPC SDK, Intel, LLVM, GNU), scientific/math libraries, and environment management tools (Environment Modules, Spack, or similar) to support and troubleshoot the HPC software stack.
- Parallel Programming Depth: Apply advanced knowledge of parallel programming models (MPI, OpenMP, threading, CUDA/HIP) to help users decompose problems, port applications, and resolve scaling issues.
- Workflow & Container Support: Support advanced computational workflows, including containerized applications and container runtimes (Charliecloud, Singularity/Apptainer, Docker, and Podman), across HPC and hybrid environments.
- Technical Representation: Represent LANL HPC Consulting at workshops, conferences, and technical meetings across the DOE Complex, sharing solutions and gathering emerging best practices.
- Knowledge Transfer: Enhance the technical expertise of junior staff through informal mentoring and knowledge-sharing, focused on real problem-solving rather than formal supervision.
Requirements
- Communication skills - demonstrated effective communication in classroom or team situations, such as technical demonstrations, system documentation, and user manuals.
- Experience with Linux-based systems.
- Demonstrated experience with HPC environments and one or more domains within them, such as parallel software, operating systems, parallel file systems, archives, parallel applications, and job schedulers.
- Programming experience (2+ years) in languages such as C, C++, Fortran, and Python.
- Proficiency in parallel programming in a CPU and/or GPU computing environment (MPI, OpenMP, CUDA, etc.)., * Advanced-level programming experience (3-5 years) in Python, or another high-level language, in addition to compiled languages (C, C++, Fortran).
- Knowledge and hands-on experience with HPC systems, including parallel filesystems, job schedulers, and interconnects.
- Experience working with or supporting scientific computing and mathematics libraries (e.g., BLAS, LAPACK, FFTW, PETSc).
- Experience with multiple Linux compilers, including NVIDIA HPC SDK, Intel, LLVM, and GNU toolchains.
- Advanced experience programming in a parallel computing environment using MPI, threading models, or both.
- Demonstrated experience with tools and methods for optimization and debugging in highly parallel environments (e.g., profilers, debuggers, tracing tools).
- Experience with scientific visualization software and tools.
- Experience using containers and container runtime technologies (Charliecloud, Singularity/Apptainer, Docker, and Podman) in HPC contexts.
Education/Experience for HPC Engineer 2: Position requires a bachelor’s in Computer Science, Computer Engineering, or a related field, and 3 years of relevant experience in high performance computing, scalable AI computing, or data center environments, or an equivalent combination of education and experience in a related field.
Education/Experience for HPC Engineer 3: Position requires a bachelor’s in Computer Science, Computer Engineering, or a related field, and 6 years of relevant experience in high performance computing, scalable AI computing, or data center environments, or an equivalent combination of education and experience in a related field.
Desired Qualifications:
- Familiarity with GPU computing in scientific environments using CUDA.
- Experience building and optimizing scientific applications at scale.
- Recent customer service experience, including use of ticketing systems.
- Experience with DevOps tools, including container runtimes and CI/CD tooling.
- Experience with code profiling, tuning, and performance optimization.
- Familiarity with AI/ML frameworks (PyTorch, TensorFlow, SciPy, NLTK).
Work Location: The work location for this position is hybrid and is located in Los Alamos. Hybrid is defined as working partially onsite/partially offsite but within 2 hours ground commute of this location. All work locations are at the discretion of management and can change at any time with appropriate notice. Current departmental policy requires a minimum of 60% on-site hours., *Eligibility requirements: To obtain a clearance, an individual must be at least 18 years of age; U.S. citizenship is required except in very limited circumstances. See DOE Order 472.2 (https://www.directives.doe.gov/directives-documents/400-series/0472.2-border-a-chg1-ltdchg/@@images/file) for additional information.
Benefits & conditions
Due to federal restrictions contained in the current National Defense Authorization Act, citizens of the People’s Republic of China-including the special administrative regions of Hong Kong and Macau-as well as citizens of the Islamic Republic of Iran, the Democratic People’s Republic of Korea (North Korea), and the Russian Federation, who are not Lawful Permanent Residents (“green card” holders) are prohibited from accessing facilities that support the mission, functions, and operations of national security laboratories and nuclear weapons production facilities, which includes Los Alamos National Laboratory.
Where You Will Work
Located in beautiful northern New Mexico, Los Alamos National Laboratory (LANL) is a multidisciplinary research institution engaged in strategic science on behalf of national security. Our generous benefits package includes:
§ PPO or High Deductible medical insurance with the same large nationwide network
§ Dental and vision insurance
§ Free basic life and disability insurance
§ Paid childbirth and parental leave
§ Award-winning 401(k) (6% matching plus 3.5% annually)
§ Learning opportunities and tuition assistance
§ Flexible schedules and time off (PTO and holidays)
§ Onsite gyms and wellness programs
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
7 Cloud Computing Trends Coming in 2025 for Developers
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
How to Become an AI Engineer