Senior Software Engineer, C++ and CUDA - Analytics...

NVIDIA Ltd.
Santa Clara, CA, United States
13 days ago
Apply on www.juju.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$184,000.0 - $287,500.0
Working hours
Regular working hours
Job source

Tech stack

Data Analysis C++ (Programming Language) Communications Protocols Nvidia CUDA Computer Programming Databases Extract Transform Load (ETL) Software Debugging Data Intelligence Open Source Technology Software Engineering Supercomputing
+8 more
Data Processing Graphics Processing Unit (GPU) Large Language Models Parallel Computation Information Technology Free and Open-Source Software Presto C++14

Job description

As part of NVIDIA’s Analytics and Data Intelligence (ADI) group, this team develops “libcudf”-the open-source CUDA C++ library that accelerates database and DataFrame operations. With flexible I/O, blazing-fast merging, aggregating, and filtering, our library serves diverse domains, including business intelligence, genomics, LLM training, and more. We use the latest tools in modern C++ and CUDA to produce software with elegant design, broad feature coverage, and best-in-class performance.

The team is looking for an outstanding engineer and/or scientist to apply their parallel programming skills to accelerate open-source software libraries for GPU-based data processing. In this position, you will drive speed-of-light performance in structured data processing, spanning hardware from single workstations to multi-node GPU supercomputers. In addition, you will be building the computational core for DataFrame and database accelerators-highly optimized C++ and CUDA libraries that leverage the parallel nature of GPUs to accelerate operations from data loading and parsing, joins, aggregations, and more. Come bring your inspiration and problem-solving skills to our open-source software suite, and you can be our next major contributor!

What you’ll be doing:

  • Own development for “UcxExchange” in Velox (30%)

  • GPU-to-GPU communication in Presto: https://github.com/facebookincubator/velox/tree/main/velox/experimental/ucx-exchange

  • Drive the feature roadmap, improve performance, and ensure correctness

  • Optimize multi-node performance for analytical queries with Presto GPU (30%)

Requirements

  • 8+ years of experience in Computer Science or Software Engineering

  • MS degree or PhD in computer science, engineering, or a related field, or equivalent experience

  • Strong Modern C++ programming skills

  • You care deeply about robust, readable, high-performance code

Ways to stand out from the crowd:

  • Expertise in high-performance communication protocols (e.g., UCX, NCCL) and distributed algorithms for query engines

  • Familiarity with RAPIDS cuDF (https://github.com/rapidsai/cudf)

  • Experience in distributed workflow development and debugging

  • Passion for open-source software development and publishing your work in technical blogs and conferences

Benefits & conditions

With competitive salaries and a generous benefits package, NVIDIA is widely considered to be one of the technology industry’s most desirable employers. We have some of the most forward-thinking and versatile people in the world working with us, and our engineering teams are growing fast in some of the most impactful fields of our generation: Analytic Systems and Data Science. If you’re an inspired Engineer or Scientist who enjoys autonomy and shares our passion for technology, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits (https://www.nvidia.com/en-us/benefits/) .

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg ¡ LIVE

1:37 min

Accelerating compute with focused developer tools

Julia Koch Julia Koch +1 ¡ World Congress 2026 Europe

4:01 min

Managing application isolation via pluggable database models

Wei Hu Wei Hu ¡ World Congress 2022

2:15 min

Introduction to CUDA and general-purpose GPU computing

Paul Graham Paul Graham ¡ World Congress 2026 Europe

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 ¡ World Congress 2026 Europe

1:52 min

Evolution of distributed SQL database management systems

Akmal Chaudhri Akmal Chaudhri ¡ LIVE

Videos

See all

Related articles

See all