Software Engineer - Work From Home

TRENT HOEK OUTDOORS LLC
Santa Clara, CA, United States
22 days ago
Apply on arc.dev
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$272,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence C++ (Programming Language) Cloud Storage Program Optimization Profiling Computer Programming Computer Networks Distributed Systems Dynamic Random-Access Memory Memory Management Python (Programming Language) Machine Learning
+13 more
Open Source Technology Peer-To-Peer (P2P) Remote Direct Memory Access Remote Access Technology Distributed Caching Data Streaming Graphics Processing Unit (GPU) Large Language Models Generative AI Low Latency Machine Learning Operations TensorRT Nvme

Job description

NVIDIA is seeking a Principal Software Engineer to define the vision and technical roadmap for memory management within large-scale LLM inference and storage systems.

The position will work closely with NVIDIA Dynamo, a high-throughput, low-latency inference framework designed for serving generative AI and reasoning models across multi-node distributed environments. Dynamo uses Rust for performance and Python for extensibility and coordinates GPU shards, request routing, and shared KV-cache management across heterogeneous clusters.

As LLM workloads increasingly exceed the memory capacity of individual GPUs, the role will focus on building infrastructure that efficiently manages data across multiple memory and storage tiers., * Design and evolve a unified memory layer spanning GPU memory, pinned host memory, RDMA-accessible memory, SSDs, and remote file, object, or cloud storage.

  • Develop systems that support large-scale LLM inference with high throughput and low latency.
  • Architect integrations with LLM serving engines such as vLLM, SGLang, and TensorRT-LLM.
  • Develop solutions for KV-cache offloading, reuse, sharing, and remote access.
  • Design interfaces and protocols supporting disaggregated prefill and peer-to-peer KV-cache sharing.
  • Build multi-tier KV-cache storage across GPU memory, CPU memory, local disks, and remote memory.
  • Work with GPU architecture, networking, and platform teams on GPUDirect, RDMA, NVLink, and related technologies.
  • Optimize KV-cache access and sharing across heterogeneous and disaggregated accelerator environments.
  • Profile and optimize systems across CPU, GPU, memory, and networking components.
  • Use performance metrics to guide architectural decisions and validate improvements in time-to-first-token (TTFT) and throughput.
  • Mentor senior and junior engineers and establish technical direction for memory and storage subsystems.
  • Lead cross-functional technical initiatives involving research, product, platform, and customer teams.
  • Represent the team in internal technical reviews, open-source communities, conferences, and customer-facing technical discussions.

Requirements

  • Master’s degree, PhD, or equivalent professional experience.
  • 15+ years of experience building large-scale distributed systems, high-performance storage, or ML infrastructure.
  • Strong programming experience with C/C++ and Python.
  • Demonstrated experience delivering production-scale services.
  • Deep knowledge of memory hierarchies, including GPU HBM, host DRAM, SSD, and remote or object storage.
  • Experience designing multi-tier systems optimized for performance and cost efficiency.
  • Experience with distributed caching or key-value systems designed for low latency and high concurrency.
  • Hands-on experience with networked I/O and technologies such as RDMA, NVMe-oF, or NVLink.
  • Understanding of disaggregated and aggregated architectures for AI clusters.
  • Strong systems profiling and optimization skills across CPU, GPU, memory, and network resources.
  • Ability to use quantitative metrics to evaluate system performance and architectural improvements.
  • Excellent communication and technical leadership skills.
  • Experience leading initiatives involving research, product, engineering, and customer-facing teams.

Preferred Qualifications

Candidates may stand out with experience in:

  • Open-source LLM serving or systems projects.
  • KV-cache optimization, compression, streaming, or reuse.
  • Unified memory or storage layers spanning GPU, host, SSD, and cloud storage.
  • Enterprise-scale or hyperscale infrastructure.
  • Memory-disaggregated architectures.
  • RDMA- or NVLink-based data planes.
  • KV-cache or CDN-style systems for machine learning.
  • Research publications or patents involving LLM systems, distributed memory, storage, or high-performance networking.

Compensation and Benefits

Benefits & conditions

The stated base salary range for this position is $272,000 to $425,500 USD, with the actual base salary determined according to factors such as location, experience, and compensation for comparable positions. The position also includes eligibility for equity and a comprehensive benefits package.

About the company

NVIDIA is a technology company known for its work in accelerated computing, artificial intelligence, GPUs, and high-performance computing. Its engineering teams develop infrastructure that supports demanding AI and machine learning workloads at large scale.

The company is expanding teams focused on generative AI infrastructure and systems engineering. This role offers an opportunity to work on technologies supporting large language model inference across distributed, heterogeneous computing environments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on arc.dev
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · World Congress 2025

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · World Congress 2026 Europe

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

2:32 min

Core libraries driving inference engines and multi-GPU networking

Adolf Hohl Adolf Hohl · World Congress 2024

2:17 min

Comparing code profiling with surface level monitoring

Jérôme Vieilledent · LIVE

Videos

See all

Related articles

See all