> Markdown version of [/jobs/ext/2661195-performance-benchmark-engineer-nvidia-gpu-systems](https://www.wearedevelopers.com/jobs/ext/2661195-performance-benchmark-engineer-nvidia-gpu-systems). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Performance/ Benchmark Engineer - NVIDIA GPU Systems - **Company:** Yoh Services LLC - **Location:** Santa Clara, CA, United States - **Salary:** $250,000.0 - $300,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Computer Clusters, Nvidia CUDA, Distributed Computing Environment, Ethernet, InfiniBand, Python (Programming Language), Machine Learning, Performance Tuning, Remote Direct Memory Access, Graphics Processing Unit (GPU), Pytorch, Large Language Models, Model Validation, Low Latency, Performance Monitor, TensorRT - **Published:** August 11, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=eb6644f0bc5ed729 ## About the Role * Deep hands-on experience with NVIDIA GPU compute platforms and AI/ML performance benchmarking. * Strong understanding of AI inference, model performance, workload characterization, and GPU architecture. * Experience with NVIDIA DGX, B200/B300, H100/H200, Blackwell, Hopper, or comparable GPU systems. * Experience analyzing performance metrics including latency, throughput, GPU utilization, memory bandwidth, and multi-GPU scaling. * Strong scripting and automation skills using Python or similar languages. Preferred Qualifications * Experience with MLPerf, CUDA, NCCL, TensorRT, Triton Inference Server, PyTorch, Nsight, or similar AI performance and profiling technologies. * Experience benchmarking LLMs, inference workloads, distributed training, or large-scale GPU clusters. * Familiarity with RDMA, RoCE, InfiniBand, Ethernet, GPUDirect RDMA, or networking considerations affecting GPU cluster performance. ## Description * Develop and execute performance benchmarks for AI inference and machine learning workloads across NVIDIA GPU systems. * Characterize performance on platforms including NVIDIA DGX and B200/B300-based systems, analyzing throughput, latency, utilization, memory behavior, and scaling efficiency. * Evaluate AI models and workload configurations to identify performance bottlenecks and recommend system or architecture improvements. * Build benchmarking methodologies, automation, and reporting frameworks to produce repeatable performance results. * Collaborate with architecture, compute, networking, and software teams to optimize end-to-end AI cluster performance. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Trends, Challenges and Best Practices for AI at the Edge](https://www.wearedevelopers.com/videos/630-trends-challenges-and-best-practices-for-ai-at-the-edge) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What’s the latest in NVIDIA CUDA Python](https://www.wearedevelopers.com/magazine/568-what-s-the-latest-in-nvidia-cuda-python) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)