Senior Reliability Engineer

NVIDIA Corporation
Santa Clara, CA, United States
1 day ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$116,000.0 - $184,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Business Analytics Applications Data Analysis Computer Programming Databases Software Debugging Hardware Design Python (Programming Language) MATLAB JMP (Statistical Software) Test Scripts

Job description

We’re seeking an outstanding Senior HTOL Reliability Engineer to join our Santa Clara lab. This role requires deep device-circuitry knowledge and hands-on hardware development. You will build next-generation HTOL boards and run HTOL processes on advanced ovens. This ensures world-class reliability of the silicon powering the AI era.

What you’ll be doing:

  • Implement and optimize HTOL test programs aligned with JEDEC standards.
  • Operate and maintain HTOL ovens, ensuring efficient test conditions and high data accuracy.
  • Design, debug, and bring up HTOL burn in boards. Debug and bring up HTOL patterns.
  • Apply sophisticated thermal management techniques to deliver detailed temperature control and mitigate thermal stress in HTOL environments.
  • Work alongside lab technicians and reliability engineers to solve technical challenges and continuously improving test processes.
  • Contribute to cross-functional teams to debug and resolve hardware and software product issues in the HTOL environment.
  • Maintain and improve our reliability database, finding opportunities for improvement.
  • Collaborate with vendors to develop and implement improvements to burn-in boards, HTOL systems, and thermal interface materials.

Requirements

  • Master’s or Bachelor’s degree in Electrical Engineering or a related field (or equivalent experience).
  • 5+ years of experience in HTOL test system operation and data analysis for semiconductor devices.
  • Proven expertise in HTOL stress testing, JEDEC standards, and environmental stress tests, including Temperature Cycling (TC), Reflow, Thermal Shock, and HAST.
  • Strong ATE or TE skills, including pattern bring up and pcb board debugging and design.
  • Hands-on experience with High power HTOL chambers including operation, repair, and preventative maintenance of HTOL chamber.
  • Proficiency with oscilloscopes, current probes, and other test equipment for data acquisition and analysis.
  • Experience with pattern vector debugging, test script development/modification, and data analysis tools.
  • Programming experience with Python or MATLAB for data analysis and automation.
  • Excellent communication, teamwork, and problem solving skills, with strong attention to detail.

Ways to stand out from the crowd:

  • Experience with multi-die HTOL testing and the associated testing challenges.
  • Background in HTOL board design for high power GPU or SoC devices.
  • Pattern translations, debugging, bringup, experience working with DFT.
  • Familiarity with reliability analytics platforms (e.g., JMP) and statistical lifetime modeling (e.g., Weibull, Arrhenius).
  • Track record of driving vendor qualification and component selection for reliability test hardware.

Benefits & conditions

With highly competitive salaries and a comprehensive benefits package, NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people on the planet working for us. If you’re creative and autonomous, with a genuine passion for technology, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 116,000 USD - 184,000 USD.

You will also be eligible for equity and benefits.

About the company

NVIDIA is the world leader in accelerated computing, developing breakthroughs that tackle challenges no one else can solve. Our work in AI and digital twins is transforming the world’s largest industries and profoundly impacting society. Come join the team and help build the next era of computing!

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

40 sec

Hardware durability labs and robot testing methods

Chris Heilmann +1 · LIVE

1:31 min

Managing self-healed test scripts safely within continuous integration

Andrei Nutas Andrei Nutas · World Congress 2026 Europe

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg · LIVE

2:40 min

Motivations for transitioning legacy MATLAB repositories to Python

Michael Niebisch Michael Niebisch · World Congress 2024

1:34 min

Profiling and debugging GPU code with Nsight developer tools

Paul Graham Paul Graham · World Congress 2025

2:04 min

Evaluating non-deterministic AI models with offline testing

Julia Kasper · Coffee With Developers

Videos

See all

Related articles

See all