Staff HPC Applications Engineer

Next Silicon Inc.
Austin, TX, United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence C++ (Programming Language) Computational Fluid Dynamics Nvidia CUDA Fortran (Programming Language) Data Flow Control OpenMP Tensorflow Software Systems Parallel Computation Information Technology

Job description

  • Profile and identify bottlenecks of a wide range of HPC applications on different architectures
  • Develop creative algorithmic and software solutions to solve bottlenecks and accelerate applications on a novel dataflow architecture
  • Translate application requirements into software features and hardware requirements
  • Develop performance models for future hardware architectures

Requirements

We are looking for an experienced Staff Applications Engineer to join the Applications team, with demonstrated ability in moving HPC applications onto new platforms and in evaluating performance on node and at scale.

Location: Hybrid in either our Austin, TX or Minneapolis, MN offices (or willing to relocate) preferred but Remote considered for exceptional candidates.

As part of a software-defined hardware company, you will play a pivotal role in driving new software features based on your analysis of customer-defined applications and measuring the resulting performance improvements. This role stands on the edge of computer science/architecture and scientific applications which span a wide range of fields, including but not limited to graph algorithms, sparse computations, weather prediction, seismic imaging, genomics, molecular dynamics, quantum chemistry, and computational fluid dynamics. If you have a passion for science, a knack for solving complex problems, and thrive in a bleeding-edge multidisciplinary environment, we want to hear from you! Requirements:

  • US citizenship and eligibility to visit US government research facilities
  • B.S. degree in a hard science, engineering, computer science, or a related field; M.S. or Ph.D. strongly preferred

  • Hands-on experience with development applications in one or more HPC domains, particularly scientific applications that run at rack or system scale
  • High level of proficiency in one of C/C++/Fortran and familiarity with the others
  • Extensive experience with node level and distributed parallel programming models and a working understanding of OpenMP, MPI in particular
  • Ability to measure application-level performance and profile HPC applications at the node level and at scale
  • Willingness to travel as necessary
  • Ability to work remotely and independently in a fast-paced environment with minimal direct supervision

Desired Skills:

  • Expertise in competitive performance analysis is strongly preferred
  • CUDA or similar GPU kernel languages
  • Hands-on experience with one or more AI/ML frameworks, such as pytorch

About the company

NextSilicon is reimagining high-performance computing. Our accelerated compute solutions leverage intelligent adaptive algorithms to vastly accelerate supercomputers, driving them forward into a new generation. Our new software-defined hardware architecture enables HPC to fulfill its promise of breakthroughs in all fields of advanced research.

At NextSilicon, everything we do is guided by three core values:

  • Professionalism: We strive for exceptional results through professionalism and unwavering dedication to quality and performance.
  • Unity: Collaboration is key to success. That’s why we foster a work environment where every employee can feel valued and heard.
  • Impact: We’re passionate about developing technologies that make a meaningful impact on industries, communities, and individuals worldwide.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:37 min

Simplifying parallel programming with the CUDA ecosystem

Paul Graham Paul Graham · LIVE

1:39 min

Fundamentals of tensors and the TensorFlow library

Håkan Silfvernagel · LIVE

6:21 min

Previewing upcoming hardware acceleration capabilities for Python environments

Chris Heilmann +2 · LIVE

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

1:37 min

Accelerating compute with focused developer tools

Julia Koch Julia Koch +1 · World Congress 2026 Europe

3:30 min

Transitioning from CUDA software architect to user

Stephen Jones · Coffee With Developers

Videos

See all

Related articles

See all