> Markdown version of [/jobs/ext/2732619-software-engineer-work-from-home](https://www.wearedevelopers.com/jobs/ext/2732619-software-engineer-work-from-home). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer - Work From Home - **Company:** TRENT HOEK OUTDOORS LLC - **Location:** Santa Clara, CA, United States (Remote available) - **Salary:** $272,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Cloud Storage, Program Optimization, Profiling, Computer Programming, Computer Networks, Distributed Systems, Dynamic Random-Access Memory, Memory Management, Python (Programming Language), Machine Learning, Open Source Technology, Peer-To-Peer (P2P), Remote Direct Memory Access, Remote Access Technology, Distributed Caching, Data Streaming, Graphics Processing Unit (GPU), Large Language Models, Generative AI, Low Latency, Machine Learning Operations, TensorRT, Nvme - **Published:** September 5, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/pho7asrggx ## About the Role * Master's degree, PhD, or equivalent professional experience. * 15+ years of experience building large-scale distributed systems, high-performance storage, or ML infrastructure. * Strong programming experience with C/C++ and Python. * Demonstrated experience delivering production-scale services. * Deep knowledge of memory hierarchies, including GPU HBM, host DRAM, SSD, and remote or object storage. * Experience designing multi-tier systems optimized for performance and cost efficiency. * Experience with distributed caching or key-value systems designed for low latency and high concurrency. * Hands-on experience with networked I/O and technologies such as RDMA, NVMe-oF, or NVLink. * Understanding of disaggregated and aggregated architectures for AI clusters. * Strong systems profiling and optimization skills across CPU, GPU, memory, and network resources. * Ability to use quantitative metrics to evaluate system performance and architectural improvements. * Excellent communication and technical leadership skills. * Experience leading initiatives involving research, product, engineering, and customer-facing teams. ### Preferred Qualifications Candidates may stand out with experience in: * Open-source LLM serving or systems projects. * KV-cache optimization, compression, streaming, or reuse. * Unified memory or storage layers spanning GPU, host, SSD, and cloud storage. * Enterprise-scale or hyperscale infrastructure. * Memory-disaggregated architectures. * RDMA- or NVLink-based data planes. * KV-cache or CDN-style systems for machine learning. * Research publications or patents involving LLM systems, distributed memory, storage, or high-performance networking. ### Compensation and Benefits ## Description NVIDIA is seeking a Principal Software Engineer to define the vision and technical roadmap for memory management within large-scale LLM inference and storage systems. The position will work closely with NVIDIA Dynamo, a high-throughput, low-latency inference framework designed for serving generative AI and reasoning models across multi-node distributed environments. Dynamo uses Rust for performance and Python for extensibility and coordinates GPU shards, request routing, and shared KV-cache management across heterogeneous clusters. As LLM workloads increasingly exceed the memory capacity of individual GPUs, the role will focus on building infrastructure that efficiently manages data across multiple memory and storage tiers., * Design and evolve a unified memory layer spanning GPU memory, pinned host memory, RDMA-accessible memory, SSDs, and remote file, object, or cloud storage. * Develop systems that support large-scale LLM inference with high throughput and low latency. * Architect integrations with LLM serving engines such as vLLM, SGLang, and TensorRT-LLM. * Develop solutions for KV-cache offloading, reuse, sharing, and remote access. * Design interfaces and protocols supporting disaggregated prefill and peer-to-peer KV-cache sharing. * Build multi-tier KV-cache storage across GPU memory, CPU memory, local disks, and remote memory. * Work with GPU architecture, networking, and platform teams on GPUDirect, RDMA, NVLink, and related technologies. * Optimize KV-cache access and sharing across heterogeneous and disaggregated accelerator environments. * Profile and optimize systems across CPU, GPU, memory, and networking components. * Use performance metrics to guide architectural decisions and validate improvements in time-to-first-token (TTFT) and throughput. * Mentor senior and junior engineers and establish technical direction for memory and storage subsystems. * Lead cross-functional technical initiatives involving research, product, platform, and customer teams. * Represent the team in internal technical reviews, open-source communities, conferences, and customer-facing technical discussions. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Your next 10x engineer isn't in your city. Refactor accordingly.](https://www.wearedevelopers.com/videos/100103-your-next-10x-engineer-isn-t-in-your-city-refactor-accordingly) - [Enhancing Workload Security in Kubernetes](https://www.wearedevelopers.com/videos/356-enhancing-workload-security-in-kubernetes) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How Much Does a Software Engineer Make? Realistic Software Engineering Salaries](https://www.wearedevelopers.com/magazine/425-how-much-does-a-software-engineer-make-realistic-software-engineering-salaries)