> Markdown version of [/jobs/ext/2727932-systems-performance-engineer](https://www.wearedevelopers.com/jobs/ext/2727932-systems-performance-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Systems Performance Engineer - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, United States - **Experience:** Expert - **Salary:** $136,000.0 - $212,750.0 - **Contract:** Permanent contract - **Skills:** Nvidia CUDA, Computer Programming, Software Debugging, General-Purpose Computing on Graphics Processing Units, Python (Programming Language), Large Language Models, Slurm, TensorRT - **Published:** September 5, 2026 - **Apply:** https://startup.jobs/senior-systems-performance-engineer-2100-nvidia-usa-9920039 ## About the Role * Ability to work on site in hardware lab environment 5 days a week * BSEE or BSCE or equivalent experience * 5+ years or more of experience in validating and debugging complex systems. * Developing/running real world ML/LLM workload. * Dynamo, TensorRT, Slurm, BCM skills mandatorily required. * Knowledge of vLLM, SG Lang preferred. * Proficiency in Cuda, Cublas and Cutlass * Deep understanding of computing architectures. * Coding experience with python programming, running simulators. * Experience with datacenter products including system management, security, networking, and storage. Ways to stand out from the crowd: * Background with x86/Arm server architectures and accelerated GPU computing. * Track record of continuous process improvement with a passion for tools and automation. ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) ## Related Articles - [What’s the latest in NVIDIA CUDA Python](https://www.wearedevelopers.com/magazine/568-what-s-the-latest-in-nvidia-cuda-python) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers)