Senior HPC Performance Engineer

NVIDIA Ltd.
Remote, OR, United States
23 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$184,000.0 - $287,500.0
Working hours
Regular working hours
Job source

Tech stack

Assembly Language C++ (Programming Language) Nvidia CUDA Computer Programming Software Debugging Fortran (Programming Language) OpenMP Software Engineering Parallel Computation Information Technology

Requirements

  • BS/MS or equivalent experience in Computer Science or related engineering field.

  • 8+ Years of programming experience.

  • Solid understanding of Fortran/C/C++, as well as programming techniques, especially for parallel architectures; preferably for compilers

  • Experience with OpenACC, OpenMP, MPI, and CUDA.

  • Strong skills in performance analysis and tuning, as well as a broad understanding of parallel applications development tools and runtime environments.

  • Strong mathematical fundamentals, including linear algebra and numerical methods.

  • Understand performance considerations, tradeoffs and impact.

  • Expert interpersonal skills, logical approach to problem solving, good time management and task prioritization skills. Excellent written and verbal communication skills, along with the ability to work in a dynamic product oriented team.

Ways to stand out from the crowd:

  • You have a deep understanding of machine architectures and micro-architectures.

  • Experience with debugging and porting as well as assembly language programming is a significant advantage.

  • Experience is leading and/or managing projects is a plus.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD.

About the company

As a member of our team in NVIDIA’s NVHPC compilers & tools group, you will analyze and run High Performance Computing (HPC) applications on HPC servers and systems to gain insight into the performance characteristics of these applications. The applications you’ll work with range from small synthetic benchmarks that use a single core to full applications that utilize all of the resources on distributed-memory systems with heterogeneous compute nodes including CPUs, GPUs and many-core processors. In this role you will analyze these applications and identify optimization opportunities for compiler development teams and application engineering teams.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:37 min

Simplifying parallel programming with the CUDA ecosystem

Paul Graham Paul Graham · LIVE

6:21 min

Previewing upcoming hardware acceleration capabilities for Python environments

Chris Heilmann +2 · LIVE

1:30 min

Drawing parallels between early compilers and AI tools

Sven Reinck Sven Reinck · Europe 2026 Virtual

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

1:37 min

Accelerating compute with focused developer tools

Julia Koch Julia Koch +1 · WWC Europe 2026

3:30 min

Transitioning from CUDA software architect to user

Stephen Jones · Coffee With Developers

Videos

See all

Related articles

See all