Senior AI and Quant DevTech Engineer

NVIDIA Corporation
Santa Clara, CA, United States
4 days ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$152,000.0 - $241,500.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence C++ (Programming Language) Profiling Nvidia CUDA Computer Programming Computer Engineering Machine Learning Monte Carlo Methods Performance Tuning Software Engineering System Software Graphics Processing Unit (GPU)
+6 more
High Performance Computing Electrical and Computer Engineering Large Language Models Parallel Computation Information Technology TensorRT

Job description

We are looking for a software engineer with a strong background in parallel processing and GPU architecture to push the limits of performance at the intersection of AI, high-performance computing, and financial markets. In this role, you will dive deep into parallel algorithms, GPUs, and sophisticated systems, identifying and eliminating bottlenecks to unlock the full power of the world’s most advanced processing hardware.

You will collaborate with top experts across industry and academia, influence next-generation platforms, and share your insights with the global developer community. Do you enjoy solving hard technical problems, love performance tuning, and want your work to have a visible impact across an entire industry? If so, we’d love for you to consider this role.

What you will be doing:

  • Designing and developing groundbreaking techniques to accelerate high-performance workloads at the intersection of AI, math, and financial systems.
  • Working hands-on with leading technical experts to analyze, optimize, and scale complex AI and HPC workloads for modern CPU and GPU architectures.
  • Profiling and eliminating performance bottlenecks across the stack-from algorithms to kernels to system-level behavior.
  • Publishing and presenting your work in conferences, talks, and blogs to educate and inspire the broader developer community.
  • Influencing the design of future hardware architectures, system software, libraries, and programming models by collaborating closely with NVIDIA research, hardware, compiler, and tools teams.

Requirements

  • Strong hands-on experience with CUDA and parallel programming.
  • Deep understanding of CPU/GPU architecture fundamentals and how they impact performance.
  • A Master’s or PhD in Computer Science, Computer Engineering, Electrical and Computer Engineering, or a related field.
  • Fluency in C/C++ and a solid foundation in algorithms and software design.
  • 5+ years of relevant work or research experience.
  • Proven experience improving the performance of large-scale computational applications on GPUs.
  • Excellent understanding of linear algebra.
  • Strong communication and organizational skills, with a logical approach to problem-solving and solid prioritization abilities.

Ways to stand out from the crowd:

  • Experience with inference optimization techniques and deploying optimized AI models in production.
  • Experience with TensorRT, TensorRT-LLM, and cuTile.
  • Experience parallelizing and optimizing machine learning methods such as decision trees, time-series models, and Monte Carlo simulations.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · World Congress 2025

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · World Congress 2026 Europe

6:21 min

Previewing upcoming hardware acceleration capabilities for Python environments

Chris Heilmann +2 · LIVE

1:37 min

Accelerating compute with focused developer tools

Julia Koch Julia Koch +1 · World Congress 2026 Europe

3:30 min

Transitioning from CUDA software architect to user

Stephen Jones · Coffee With Developers

Videos

See all

Related articles

See all