CUDA/C++ Performance Engineer for Differentiable Physics Simulator

HONDA RESEARCH INSTITUTE USA
San Jose, CA, United States
2 months ago
Apply on honda-ri.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience required
1 year minimum
Working hours
Regular working hours
Job source

Tech stack

C++ (Programming Language) Computer Simulation Nvidia CUDA Data Streaming

Job description

  • Profile bottlenecks across CUDA kernels, memory transfers, synchronization, sparse assembly, solver steps, and differentiable rollout paths.
  • Determine the highest-impact performance lever for each bottleneck, whether kernel tuning, data residency, batching, stream usage, solver changes, or reduction/assembly redesign.
  • Improve existing CUDA backend architecture, including host/device data flow, CUDA kernels, sparse assembly, and solver structure.
  • Evaluate tradeoffs between targeted optimization, architectural refactoring, and larger rewrites when justified by profiling evidence.
  • Measure and validate improvements in speed, GPU utilization, correctness, and numerical reproducibility.
  • Deliver results in accordance with project timelines.
  • Prepare written and oral technical reports and demonstrations.
  • Collaborate with our teams of scientists and engineers in Honda’s regional and global R&D offices. Communicate profiling results, tradeoffs, and implementation outcomes to audiences with varying CUDA experience.

Requirements

Honda Research Institute USA (HRI-US) is seeking a self-motivated engineer to join our Intelligent Robotics Research division. This individual will improve performance of a CUDA/C++ differentiable physics simulator across the GPU backend, including CUDA kernels, host/device data flow, sparse solver structure, and backward workflows used in optimization. The work will require profiling forward and backward simulation workloads, identifying bottlenecks, improving GPU utilization, and reducing CPU/GPU synchronization and transfer overhead., * Strong C++ and CUDA C++ experience in production or research codebases.

  • Proven experience profiling and optimizing CUDA kernels with tools such as NVIDIA Nsight Systems, Nsight Compute, or equivalent GPU profiling workflows.
  • Comfortable editing low-level GPU code involving reductions, atomics, sparse matrices, memory coalescing, launch configuration, and synchronization.
  • Experience reducing CPU/GPU transfer overhead using better data residency, batching, pinned memory, async copies, streams, or kernel fusion.
  • Familiarity with numerical simulation, optimization, or differentiable physics workflows.
  • At least 1 year of hands-on experience with the qualifications above.

Bonus Qualifications

  • Experience with contact-rich physics simulation.
  • Experience with differentiable simulation or trajectory optimization.
  • Experience optimizing sparse linear solvers on GPU.
  • Experience tuning CUDA MPS workloads or multi-process GPU scheduling.
  • 3+ years of hands-on experience with the qualifications above preferred.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on honda-ri.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:30 min

Transitioning from CUDA software architect to user

Stephen Jones · Coffee With Developers

6:21 min

Previewing upcoming hardware acceleration capabilities for Python environments

Chris Heilmann +2 · LIVE

1:50 min

Lowering pipeline latency with data streaming

Nathaniel Okenwa Nathaniel Okenwa · World Congress 2024

1:03 min

Memory requirements for modern quantum computing hardware simulation

Tomislav Tipurić Tomislav Tipurić · World Congress 2023

1:37 min

Accelerating compute with focused developer tools

Julia Koch Julia Koch +1 · World Congress 2026 Europe

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all