> Markdown version of [/jobs/ext/656553-cuda-c-performance-engineer-for-differentiable-physics-simulator](https://www.wearedevelopers.com/jobs/ext/656553-cuda-c-performance-engineer-for-differentiable-physics-simulator). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # CUDA/C++ Performance Engineer for Differentiable Physics Simulator - **Company:** HONDA RESEARCH INSTITUTE USA - **Location:** San Jose, CA, United States - **Contract:** Temporary contract - **Skills:** C++ (Programming Language), Computer Simulation, Nvidia CUDA, Data Streaming - **Published:** June 26, 2026 - **Apply:** https://usa.honda-ri.com/-/cuda/c-performance-engineer-for-differentiable-physics-simulator?redirect=%2Fcontract-positions ## About the Role Honda Research Institute USA (HRI-US) is seeking a self-motivated engineer to join our Intelligent Robotics Research division. This individual will improve performance of a CUDA/C++ differentiable physics simulator across the GPU backend, including CUDA kernels, host/device data flow, sparse solver structure, and backward workflows used in optimization. The work will require profiling forward and backward simulation workloads, identifying bottlenecks, improving GPU utilization, and reducing CPU/GPU synchronization and transfer overhead., * Strong C++ and CUDA C++ experience in production or research codebases. * Proven experience profiling and optimizing CUDA kernels with tools such as NVIDIA Nsight Systems, Nsight Compute, or equivalent GPU profiling workflows. * Comfortable editing low-level GPU code involving reductions, atomics, sparse matrices, memory coalescing, launch configuration, and synchronization. * Experience reducing CPU/GPU transfer overhead using better data residency, batching, pinned memory, async copies, streams, or kernel fusion. * Familiarity with numerical simulation, optimization, or differentiable physics workflows. * At least 1 year of hands-on experience with the qualifications above. Bonus Qualifications * Experience with contact-rich physics simulation. * Experience with differentiable simulation or trajectory optimization. * Experience optimizing sparse linear solvers on GPU. * Experience tuning CUDA MPS workloads or multi-process GPU scheduling. * 3+ years of hands-on experience with the qualifications above preferred. ## Description * Profile bottlenecks across CUDA kernels, memory transfers, synchronization, sparse assembly, solver steps, and differentiable rollout paths. * Determine the highest-impact performance lever for each bottleneck, whether kernel tuning, data residency, batching, stream usage, solver changes, or reduction/assembly redesign. * Improve existing CUDA backend architecture, including host/device data flow, CUDA kernels, sparse assembly, and solver structure. * Evaluate tradeoffs between targeted optimization, architectural refactoring, and larger rewrites when justified by profiling evidence. * Measure and validate improvements in speed, GPU utilization, correctness, and numerical reproducibility. * Deliver results in accordance with project timelines. * Prepare written and oral technical reports and demonstrations. * Collaborate with our teams of scientists and engineers in Honda's regional and global R&D offices. Communicate profiling results, tradeoffs, and implementation outcomes to audiences with varying CUDA experience. ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Quantum Computing: Separating the Breakthroughs from the Buzz](https://www.wearedevelopers.com/videos/100314-quantum-computing-separating-the-breakthroughs-from-the-buzz) - [Python-Based Data Streaming Pipelines Within Minutes](https://www.wearedevelopers.com/videos/1233-python-based-data-streaming-pipelines-within-minutes) - [CUDA Python: GPU programming for the modern developer](https://www.wearedevelopers.com/videos/100221-cuda-python-gpu-programming-for-the-modern-developer) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/1521-accelerating-python-on-gpus) ## Related Articles - [What’s the latest in NVIDIA CUDA Python](https://www.wearedevelopers.com/magazine/568-what-s-the-latest-in-nvidia-cuda-python) - [Top 6 Hackathons for Developers in 2023](https://www.wearedevelopers.com/magazine/263-top-6-hackathons-for-developers-in-2023) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers) - [Dev Digest 157: CUDA in Python, Gemini Code Assist and Back-dooring LLMs](https://www.wearedevelopers.com/magazine/557-dev-digest-157-cuda-in-python-gemini-code-assist-and-back-dooring-llms)