> Markdown version of [/jobs/ext/1953421-systems-performance-engineer](https://www.wearedevelopers.com/jobs/ext/1953421-systems-performance-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Systems Performance Engineer - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States - **Experience:** Expert - **Salary:** $184,000.0 - $287,500.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Cloud Computing, Profiling, Data Compression, Nvidia CUDA, Computer Programming, Computer Literacy, Data Structures, Design of User Interfaces, Multiprocessing, PCI Express, Memory Leaks, Graphics Processing Unit (GPU), Perf (Linux), Kubernetes, Information Technology, Slurm - **Published:** August 6, 2026 - **Apply:** https://www.careerbuilder.com/job-details/senior-performance-architect-heterogeneous-workload-optimization-santa-clara-ca--50b12467-205d-45f6-9272-ca8302ad94bc ## About the Role * A grasp of the CUDA programming model and experience employing GPU profiling tools like NVIDIA Nsight Systems/Compute to address PCIe bottlenecks and kernel stalls. * Extensive knowledge of profiling tools such as perf, eBPF, VTune, or Valgrind, along with insight into their internal mechanisms. * A passion for meticulous benchmarking and the ability to distill sophisticated performance data into actionable engineering roadmaps. * Experience with distributed compute environments (Slurm, LSF, or Kubernetes). * A BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field (or equivalent experience) with more than 8+yrs of relevant experience and at least 5 years involved in systems-level performance analysis., Artificial Intelligence (AI), Autonomous Driving Systems, Benchmarking, CPU (Central Processing Unit), CUDA (Compute Unified Device Architecture), Cloud Computing, Compensation and Benefits, Complexity Algorithms, Computer Science, Computer Skills, Computer Systems, Data Compression, Data Structures, Electrical Engineering, GPU (Graphics Processing Unit), High Throughput, Kernel Programming, Memory Hardware, PCI Express (PCI-E), Performance Analysis, Performance Engineering, Predictive Modeling, Purchasing/Procurement, Sockets, Systems Engineering, User Interface Design ## Description As EDA workloads transition from traditional CPU-bound tasks to massively parallel GPU-accelerated engines, the complexity of identifying bottlenecks has scaled exponentially. We are seeking a Senior Systems Performance Engineer to build our next generation of profiling infrastructure. You will be responsible for measuring, analyzing, and optimizing the interaction between extensive design graphs in system memory and high-throughput kernels on the GPU. Join us to push the boundaries of whats possible in the future of computing!, * Architecting and maintaining custom profiling frameworks that provide a unified view of execution across CPU (multi-core/multi-socket) and GPU (multi-node/NVLink) environments. * Conducting deep-dive benchmarking of EDA applications to characterize memory access patterns, cache hit rates, and instruction-level parallelism. * Using GPU profilers to detect GPU-side inefficiencies such as warp divergence, sub-optimal occupancy, and PCIe/NVLink bottlenecks. * Developing tools to monitor and attribute high-watermark memory usage in multi-terabyte EDA builds, finding opportunities for data structure compression or smarter memory pooling. * Developing predictive models to guide hardware procurement and cloud instance selection based on built gate-count and algorithmic complexity. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/1521-accelerating-python-on-gpus) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) ## Related Articles - [What’s the latest in NVIDIA CUDA Python](https://www.wearedevelopers.com/magazine/568-what-s-the-latest-in-nvidia-cuda-python) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [How software is steering vehicle technology](https://www.wearedevelopers.com/magazine/515-how-software-is-steering-vehicle-technology)