Senior Software Architect, AI Systems...

NVIDIA Ltd.
Santa Clara, CA, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$224,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon S3 C++ (Programming Language) Profiling Computer Engineering Software Debugging Distributed Computing Environment InfiniBand Storage Area Network (SAN) Remote Direct Memory Access Software Engineering System Programming Reinforcement Learning
+5 more
Large Language Models Information Technology Machine Learning Operations TensorRT Nvme

Requirements

  • 12+ years in systems software and/or networking with demonstrated ownership of complex projects.

  • MS, PhD or equivalent experience in Computer Science, Computer Engineering, Electrical Engineering, or a related field.

  • Solid understanding of high-performance networking: InfiniBand, RoCE, RDMA, NVLink, GPUDirect.

  • Strong C/C++/Rust systems programming with comfort in performance profiling and low-level debugging.

  • Understanding of ML systems concepts-transformer architectures, KV cache mechanics, model parallelism, or distributed training and inference patterns.

Ways to stand out from the crowd:

  • Knowledge of ML inference frameworks (vLLM, SGLang, TensorRT-LLM) and their communication requirements.

  • Knowledge of storage networking (NVMe-oF, GPUDirect Storage, S3).

  • Background of Reinforcement Learning systems.

Benefits & conditions

With competitive salaries and a comprehensive benefits package, NVIDIA is widely regarded as one of the most desirable technology employers in the world. Our teams are composed of some of the most forward-thinking and driven engineers in the industry, and we continue to grow rapidly. If you are a senior data engineer passionate about building large-scale, high-impact data platforms, we’d love to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 5, and 272,000 USD - 431,250 USD for Level 6.

You will also be eligible for equity and benefits (https://www.nvidia.com/en-us/benefits/) .

About the company

NVIDIA (Santa Clara, CA)

An applied research team within NVIDIA’s Networking Systems & Software Architecture group is solving some of AI’s hardest infrastructure problems. The team builds systems-level software that moves data between GPUs, nodes, and storage at the speed modern AI demands-spanning low-level transport optimization, hardware-software co-design, and communication frameworks that plug directly into production AI stacks. The team’s charter expands into emerging domains including quantum computing interconnects.

The Senior Architect role is to own modules and projects end-to-end-from scoping research questions to shipping production code. It calls for a recognized expert who drives technical decisions, pulls in ideas from research and industry, and regularly prototypes new approaches to prove a point. The work lives at the boundary of applied research and production engineering!

What you will be doing:

  • Architecting and implementing high-performance communication and memory management libraries for distributed AI

  • Driving hardware-software co-optimization with GPU, DPU, NIC, and switch teams through GPUDirect RDMA, NVLink, and next-generation interconnects

  • Profiling and optimizing data movement across GPU memory, system DRAM, NVMe, and network fabrics

  • Integrating networking capabilities into AI serving stacks such as vLLM, SGLang, and TensorRT-LLM

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · WWC 2025

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · WWC Europe 2026

2:32 min

Core libraries driving inference engines and multi-GPU networking

Adolf Hohl Adolf Hohl · WWC 2024

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

2:17 min

Comparing code profiling with surface level monitoring

Jérôme Vieilledent · LIVE

Videos

See all

Related articles

See all