Manager, System Level Reliability Test

NVIDIA Ltd.
Santa Clara, CA, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$184,000.0 - $287,500.0
Working hours
Regular working hours
Job source

Tech stack

Computer Graphics Laboratory Information Management Systems Reliability Engineering Software Reliability Testing Reliability of Systems

Job description

For more than 25 years, NVIDIA has redefined computer graphics, PC gaming, and accelerated computing. It’s a distinctive heritage of innovation that’s motivated by great technology-and outstanding people. The Reliability Test Manager owns the planning, execution, and management of reliability testing programs for NVIDIA Board and System products and related components. This role ensures compliance with industry standards and customer requirements for long-term reliability, quality, and performance. The manager drives central initiatives, coordinates test methodologies, and develops continuous improvement in reliability processes.

What you’ll be doing:

  • Laboratory and People Management: System reliability lab management, maintenance, and expansion. Supervision of lab engineering and technician personnel.

  • Strategic Leadership: Provide direction for reliability test programs aligned with interpersonal goals; mentor and develop team members.

  • Program Management: Direct and supervise execution of reliability test plans, encompassing qualification and accelerated life testing.

  • Standards Compliance: Ensure alignment with standards such as ISO, JEDEC, EIA, IEC, ISTA, NEBS, IEC, and customer-specific reliability standards.

  • Failure Analysis: Participate in root cause investigations and corrective actions for reliability-related failures.

  • Multi-functional Collaboration: Partner with various engineering groups, including process, and packaging teams to apply design-for-reliability principles.

  • Reporting: Deliver detailed reliability reports and communicate findings to internal collaborators and customers.

  • Continuous Improvement: Implement initiatives to improve test efficiency, accuracy, and lab throughput.

Requirements

  • Bachelor or master’s in electrical or mechanical engineering, or related field.

  • 8+ overall years in electronic product reliability testing including 4+ years reliability lab management.

  • Proficiency in accelerated Reliability evaluations, environmental stress testing, failure analysis, reliability modeling, and statistical analysis of electronic products; strong leadership and project management.

Ways to stand out from the crowd:

  • Experience with various environmental and mechanical test equipment.

  • Experience in crafting reliability testing scripts based on test profiles provided by customers for automotive, servers and or embedded products

  • Knowledge of reliability predictions principles.

  • Thrive in a fast moving, adventurous and highly multi-functional and multi-cultural environment.

Benefits & conditions

Recognized as one of the most sought-after employers in the technology sector, NVIDIA provides highly attractive salary packages along with an extensive range of benefits. We have some of the most forward-thinking and hardworking people in the world working for us. Do you love challenges? If so, we want to hear from you. When considering your future, explore what we can provide for you and your loved ones.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD.

You will also be eligible for equity and benefits (https://www.nvidia.com/en-us/benefits/) .

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

40 sec

Hardware durability labs and robot testing methods

Chris Heilmann +1 · LIVE

2:08 min

History and scale of NVIDIA GPU computing

Paul Graham Paul Graham · LIVE

5:26 min

Balancing continuous deployment with strict system reliability

David Singleton David Singleton +1 · Coffee With Developers

1:06 min

Empowering site reliability engineers with integrated AI agents

Osmar Matos Osmar Matos · WWC Europe 2026

1:18 min

Cultivating system reliability as a critical leadership responsibility

Robert Barron Robert Barron · WWC Europe 2026

1:38 min

Adopting site reliability engineering practices for machine learning

Cassie Kozyrkov · WWC 2022

Videos

See all

Related articles

See all