HPC Engineer 2/3

Los Alamos National Laboratory
Los Alamos, NM, United States
2 days ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$128,000.0 - $215,900.0
Working hours
Regular working hours
Job source

Tech stack

C (Programming Language) Artificial Intelligence Basic Linear Algebra Subprograms C++ (Programming Language) Compilers Software Documentation Profiling Nvidia CUDA Computer Programming Computer Engineering Data Centers Software Debugging
+28 more
Linux DevOps File Systems Memory Management Fortran (Programming Language) General-Purpose Computing on Graphics Processing Units Issue Tracking Systems Job Scheduling Python (Programming Language) Linux System Administration Machine Learning NLTK (NLP Analysis) OpenMP Package Management Systems Performance Tuning Tensorflow Scientific Computating SciPy Supercomputing User Environment Management High Performance Computing Pytorch Parallel Computation Containerization Information Technology Modeling and Simulation Engineering Base Docker

Job description

Join the High Performance Computing Environments Group (HPC-ENV) at Los Alamos National Laboratory, where we manage and operate advanced large-scale computing infrastructure supporting a diverse range of critical workloads. The HPC Consulting team provides direct support to HPC customers to enhance the productivity of the user community by providing quality technical support for a wide range of workloads, including:

  • Large-Scale Modeling & Simulation (e.g., physics, engineering, climate)
  • Artificial Intelligence & Machine Learning (AI/ML) deployments
  • Emerging software for next-generation computing needs

Team Responsibilities

Our Consultants are a user-facing team who provide support and engage with LANL customers via documentation, consulting, and training. As HPC Consultants, we are a single point of contact for HPC systems and deliver customer-focused services, tools, and software. The Division’s goal is to create an effective HPC environment in which scientists can be as productive as possible on world-class supercomputing systems.

The HPC Division supports the Los Alamos National Laboratory (LANL) mission by managing a world-class supercomputing center. We support stockpile stewardship for NNSA/DOE and accelerate scientific discovery for scientists. We integrate and support some of the world’s largest supercomputers during an exciting time in computing, with a focus on traditional large-scale simulations, data science, artificial intelligence, and machine learning.

HPC Engineer 2 ($106,400-176,000)

  • Direct Customer Support: Provide expert guidance and consultation to scientists, engineers, and other HPC users to enhance their workflows and enable effective use of HPC systems through phone, email, ticketing systems, and in-person consultations.
  • Documentation and Training: Develop user guides, best practices, and training materials to help the user community get the most from the HPC environment.
  • System Efficiency: Work with HPC system administrators to identify issues and improve overall system performance and resource utilization.
  • Independent Problem Solving: Work independently to implement customer solutions to current problems and enhance the HPC user environment.
  • External Representation: Represent LANL HPC Consulting at workshops, meetings, and conferences with other HPC and NNSA/DOE sites.
  • HPC Application Support: Provide in-depth support for users of HPC systems, including installation, configuration, and troubleshooting of scientific and engineering applications.
  • Performance Optimization: Analyze and optimize the performance of parallel applications and workflows to ensure efficient use of HPC resources.
  • Code Profiling and Tuning: Collaborate with software developers and users to profile and tune applications for performance, memory usage, and scalability on HPC platforms.

HPC Engineer 3 ($128,000-215,900)

In addition to the duties outlined above, a successful HPC Engineer 3 candidate will be required to:

  • Advanced Application Support: Serve as the go-to technical resource for the most complex scientific and engineering application issues, including deep debugging of parallel codes, build failures, and cross-platform portability problems.
  • Performance Engineering: Conduct advanced profiling, tuning, and optimization of parallel applications across CPU and GPU architectures, identifying bottlenecks in memory, I/O, communication, and scalability.
  • Software Stack Expertise: Maintain deep working knowledge of compilers (NVIDIA HPC SDK, Intel, LLVM, GNU), scientific/math libraries, and environment management tools (Environment Modules, Spack, or similar) to support and troubleshoot the HPC software stack.
  • Parallel Programming Depth: Apply advanced knowledge of parallel programming models (MPI, OpenMP, threading, CUDA/HIP) to help users decompose problems, port applications, and resolve scaling issues.
  • Workflow & Container Support: Support advanced computational workflows, including containerized applications and container runtimes (Charliecloud, Singularity/Apptainer, Docker, and Podman), across HPC and hybrid environments.
  • Technical Representation: Represent LANL HPC Consulting at workshops, conferences, and technical meetings across the DOE Complex, sharing solutions and gathering emerging best practices.
  • Knowledge Transfer: Enhance the technical expertise of junior staff through informal mentoring and knowledge-sharing, focused on real problem-solving rather than formal supervision.

Requirements

  • Communication skills - demonstrated effective communication in classroom or team situations, such as technical demonstrations, system documentation, and user manuals.
  • Experience with Linux-based systems.
  • Demonstrated experience with HPC environments and one or more domains within them, such as parallel software, operating systems, parallel file systems, archives, parallel applications, and job schedulers.
  • Programming experience (2+ years) in languages such as C, C++, Fortran, and Python.
  • Proficiency in parallel programming in a CPU and/or GPU computing environment (MPI, OpenMP, CUDA, etc.)., * Advanced-level programming experience (3-5 years) in Python, or another high-level language, in addition to compiled languages (C, C++, Fortran).
  • Knowledge and hands-on experience with HPC systems, including parallel filesystems, job schedulers, and interconnects.
  • Experience working with or supporting scientific computing and mathematics libraries (e.g., BLAS, LAPACK, FFTW, PETSc).
  • Experience with multiple Linux compilers, including NVIDIA HPC SDK, Intel, LLVM, and GNU toolchains.
  • Advanced experience programming in a parallel computing environment using MPI, threading models, or both.
  • Demonstrated experience with tools and methods for optimization and debugging in highly parallel environments (e.g., profilers, debuggers, tracing tools).
  • Experience with scientific visualization software and tools.
  • Experience using containers and container runtime technologies (Charliecloud, Singularity/Apptainer, Docker, and Podman) in HPC contexts.

Education/Experience for HPC Engineer 2: Position requires a bachelor’s in Computer Science, Computer Engineering, or a related field, and 3 years of relevant experience in high performance computing, scalable AI computing, or data center environments, or an equivalent combination of education and experience in a related field.

Education/Experience for HPC Engineer 3: Position requires a bachelor’s in Computer Science, Computer Engineering, or a related field, and 6 years of relevant experience in high performance computing, scalable AI computing, or data center environments, or an equivalent combination of education and experience in a related field.

Desired Qualifications:

  • Familiarity with GPU computing in scientific environments using CUDA.
  • Experience building and optimizing scientific applications at scale.
  • Recent customer service experience, including use of ticketing systems.
  • Experience with DevOps tools, including container runtimes and CI/CD tooling.
  • Experience with code profiling, tuning, and performance optimization.
  • Familiarity with AI/ML frameworks (PyTorch, TensorFlow, SciPy, NLTK).

Work Location: The work location for this position is hybrid and is located in Los Alamos. Hybrid is defined as working partially onsite/partially offsite but within 2 hours ground commute of this location. All work locations are at the discretion of management and can change at any time with appropriate notice. Current departmental policy requires a minimum of 60% on-site hours., *Eligibility requirements: To obtain a clearance, an individual must be at least 18 years of age; U.S. citizenship is required except in very limited circumstances. See DOE Order 472.2 (https://www.directives.doe.gov/directives-documents/400-series/0472.2-border-a-chg1-ltdchg/@@images/file) for additional information.

Benefits & conditions

Due to federal restrictions contained in the current National Defense Authorization Act, citizens of the People’s Republic of China-including the special administrative regions of Hong Kong and Macau-as well as citizens of the Islamic Republic of Iran, the Democratic People’s Republic of Korea (North Korea), and the Russian Federation, who are not Lawful Permanent Residents (“green card” holders) are prohibited from accessing facilities that support the mission, functions, and operations of national security laboratories and nuclear weapons production facilities, which includes Los Alamos National Laboratory.

Where You Will Work

Located in beautiful northern New Mexico, Los Alamos National Laboratory (LANL) is a multidisciplinary research institution engaged in strategic science on behalf of national security. Our generous benefits package includes:

§ PPO or High Deductible medical insurance with the same large nationwide network

§ Dental and vision insurance

§ Free basic life and disability insurance

§ Paid childbirth and parental leave

§ Award-winning 401(k) (6% matching plus 3.5% annually)

§ Learning opportunities and tuition assistance

§ Flexible schedules and time off (PTO and holidays)

§ Onsite gyms and wellness programs

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:37 min

Simplifying parallel programming with the CUDA ecosystem

Paul Graham Paul Graham · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

4:54 min

Development history of scientific computation libraries and PyViz tools

Radovan Kavický · LIVE

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

2:52 min

Addressing junior hiring bottlenecks and mitigating widespread employee burnout

Hung Lee Hung Lee +3 · World Congress 2026 Europe

Videos

See all

Related articles

See all