Performance/ Benchmark Engineer - NVIDIA GPU Systems

Yoh Services LLC
Santa Clara, CA, United States
25 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$250,000.0 - $300,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Computer Clusters Nvidia CUDA Distributed Computing Environment Ethernet InfiniBand Python (Programming Language) Machine Learning Performance Tuning Remote Direct Memory Access Graphics Processing Unit (GPU) Pytorch
+5 more
Large Language Models Model Validation Low Latency Performance Monitor TensorRT

Job description

  • Develop and execute performance benchmarks for AI inference and machine learning workloads across NVIDIA GPU systems.
  • Characterize performance on platforms including NVIDIA DGX and B200/B300-based systems, analyzing throughput, latency, utilization, memory behavior, and scaling efficiency.
  • Evaluate AI models and workload configurations to identify performance bottlenecks and recommend system or architecture improvements.
  • Build benchmarking methodologies, automation, and reporting frameworks to produce repeatable performance results.
  • Collaborate with architecture, compute, networking, and software teams to optimize end-to-end AI cluster performance.

Requirements

  • Deep hands-on experience with NVIDIA GPU compute platforms and AI/ML performance benchmarking.
  • Strong understanding of AI inference, model performance, workload characterization, and GPU architecture.
  • Experience with NVIDIA DGX, B200/B300, H100/H200, Blackwell, Hopper, or comparable GPU systems.
  • Experience analyzing performance metrics including latency, throughput, GPU utilization, memory bandwidth, and multi-GPU scaling.
  • Strong scripting and automation skills using Python or similar languages.

Preferred Qualifications

  • Experience with MLPerf, CUDA, NCCL, TensorRT, Triton Inference Server, PyTorch, Nsight, or similar AI performance and profiling technologies.
  • Experience benchmarking LLMs, inference workloads, distributed training, or large-scale GPU clusters.
  • Familiarity with RDMA, RoCE, InfiniBand, Ethernet, GPUDirect RDMA, or networking considerations affecting GPU cluster performance.

Benefits & conditions

Pulled from the full job description

  • Referral program
  • 401(k)
  • Health insurance
  • Vision insurance
  • Health savings account
  • Dental insurance
  • Employee assistance program, Estimated Min Rate: $250,000.00/Annually Estimated Max Rate: $300,000.00/Annually

What’s In It for You? We welcome you to be a part of the largest and legendary global staffing companies to meet your career aspirations. Yoh’s network of client companies has been employing professionals like you for over 65 years in the U.S., UK and Canada. Join Yoh’s extensive talent community that will provide you with access to Yoh’s vast network of opportunities and gain access to this exclusive opportunity available to you. Benefit eligibility is in accordance with applicable laws and client requirements. Benefits include:

  • Medical, Prescription, Dental & Vision Benefits (for employees working 20+ hours per week)
  • Health Savings Account (HSA) (for employees working 20+ hours per week)
  • Life & Disability Insurance (for employees working 20+ hours per week)
  • MetLife Voluntary Benefits
  • Employee Assistance Program (EAP)
  • 401K Retirement Savings Plan
  • Direct Deposit & weekly epayroll
  • Referral Bonus Programs
  • Certification and training opportunities

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · World Congress 2025

4:52 min

Connecting namespaces with local virtual ethernet pairs

Oliver Seitz Oliver Seitz · World Congress 2025

4:54 min

NVIDIA local and edge AI hardware capabilities overview

Joerg Krall Joerg Krall · World Congress 2026 Europe

2:32 min

Core libraries driving inference engines and multi-GPU networking

Adolf Hohl Adolf Hohl · World Congress 2024

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all