Senior Software Engineer, CUDA Core Libraries

NVIDIA Ltd.
Las Cruces, NM, United States
26 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$100,000.0 - $130,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) C++ (Programming Language) Software Documentation Profiling Code Review Nvidia CUDA Computer Programming Computer Engineering Continuous Integration Programming Tools Python (Programming Language) Performance Tuning
+10 more
Software Engineering Openapi Rust (Programming Language) Graphics Processing Unit (GPU) Pytorch Gpu Programming Information Technology Build Tools Api Design C++14

Job description

  • Design and implement idiomatic Python APIs and bindings for foundational CUDA capabilities and GPU algorithms.
  • Develop and integrate the native C/C++ components that support Python-facing functionality.
  • Define reliable and efficient interoperability boundaries between Python, C/C++, Rust, and other languages.
  • Develop high-performance interfaces that minimize Python and native-language integration overhead.
  • Own features throughout their lifecycle: design, implementation, testing, profiling, benchmarking, documentation, release, and long-term maintenance.
  • Improve the Python developer experience through typing, packaging, examples, diagnostics, continuous integration, and compatibility testing.
  • Collaborate with C/C++, Rust, compiler, and runtime engineers on shared architecture and API decisions.
  • Work directly with users to investigate correctness, usability, compatibility, and performance issues.

Requirements

We are hiring a Senior Software Engineer to advance the Python experience for CUDA Core Libraries. You will build Pythonic APIs, language bindings, algorithms, and runtime infrastructure on top of native C/C++ foundations. You will join the team building the foundational libraries, algorithms, and language/runtime infrastructure that make CUDA a speed-of-light experience for developers and AI coding agents alike., * BS, MS, or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.

  • 8+ years of relevant software-development experience.
  • Strong production programming skills in both Python and C/C++; both are required for this role.
  • Experience building Python interfaces to native or systems-level software.
  • Understanding of systems software concepts, performance, concurrency, and API design.
  • Practical experience with parallel, heterogeneous, or GPU programming.
  • Experience developing production software or widely used libraries, including testing, profiling, benchmarking, packaging, and code review.
  • Ability to work independently, define project scope, and drive complex work to completion.
  • Clear written communication skills for API specifications, technical designs, and user documentation.
  • Comfort working in large codebases spanning Python, C/C++, build systems, packaging, and continuous-integration infrastructure.

Ways to stand out from the crowd:

  • Strong understanding of CPU/GPU architecture and performance optimization, with hands-on experience in GPU-accelerated stacks (CUDA C++/Python, PyTorch, JAX, Numba, CuPy, or similar).
  • Proficiency with modern C++ and GPU libraries such as Thrust, CUB, and libcudacxx.
  • Experience with compiler infrastructure and tooling, including LLVM, Clang, or MLIR.
  • Expertise in designing low-overhead interoperability between Python and native languages, including exposure to Rust in mixed-language stacks.
  • Demonstrated interest in developer tools, library design, and improving developer productivity.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5., $100,000.00 per year

About the company

NVIDIA’s accelerated computing platform is foundational to modern HPC and AI. At the center of this platform are CUDA Core Libraries that enable developers to build fast, reliable, and scalable GPU-accelerated software.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jofdav.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:53 min

Using OpenAPI as an executable contract

Violina Popova Violina Popova · Europe 2026 Virtual

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

4:41 min

Replacing PyTorch with ONNX runtime for AWS Lambda deployments

Marek Suppa · LIVE

1:11 min

The expanded CUDA ecosystem and native Python support

Paul Graham Paul Graham · WWC Europe 2026

2:48 min

Understanding the purpose of OpenAPI specifications

Christopher Walles Christopher Walles · WWC 2024

Videos

See all

Related articles

See all