CUDA/C++ Performance Engineer for Differentiable Physics Simulator
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
- Profile bottlenecks across CUDA kernels, memory transfers, synchronization, sparse assembly, solver steps, and differentiable rollout paths.
- Determine the highest-impact performance lever for each bottleneck, whether kernel tuning, data residency, batching, stream usage, solver changes, or reduction/assembly redesign.
- Improve existing CUDA backend architecture, including host/device data flow, CUDA kernels, sparse assembly, and solver structure.
- Evaluate tradeoffs between targeted optimization, architectural refactoring, and larger rewrites when justified by profiling evidence.
- Measure and validate improvements in speed, GPU utilization, correctness, and numerical reproducibility.
- Deliver results in accordance with project timelines.
- Prepare written and oral technical reports and demonstrations.
- Collaborate with our teams of scientists and engineers in Honda’s regional and global R&D offices. Communicate profiling results, tradeoffs, and implementation outcomes to audiences with varying CUDA experience.
Requirements
Honda Research Institute USA (HRI-US) is seeking a self-motivated engineer to join our Intelligent Robotics Research division. This individual will improve performance of a CUDA/C++ differentiable physics simulator across the GPU backend, including CUDA kernels, host/device data flow, sparse solver structure, and backward workflows used in optimization. The work will require profiling forward and backward simulation workloads, identifying bottlenecks, improving GPU utilization, and reducing CPU/GPU synchronization and transfer overhead., * Strong C++ and CUDA C++ experience in production or research codebases.
- Proven experience profiling and optimizing CUDA kernels with tools such as NVIDIA Nsight Systems, Nsight Compute, or equivalent GPU profiling workflows.
- Comfortable editing low-level GPU code involving reductions, atomics, sparse matrices, memory coalescing, launch configuration, and synchronization.
- Experience reducing CPU/GPU transfer overhead using better data residency, batching, pinned memory, async copies, streams, or kernel fusion.
- Familiarity with numerical simulation, optimization, or differentiable physics workflows.
- At least 1 year of hands-on experience with the qualifications above.
Bonus Qualifications
- Experience with contact-rich physics simulation.
- Experience with differentiable simulation or trajectory optimization.
- Experience optimizing sparse linear solvers on GPU.
- Experience tuning CUDA MPS workloads or multi-process GPU scheduling.
- 3+ years of hands-on experience with the qualifications above preferred.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Top 6 Hackathons for Developers in 2023
How software is steering vehicle technology
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Résumé-Driven Development: How IT trends affect the job market for software developers