Senior Datacenter Technical Program Manager,...

NVIDIA Ltd.
Santa Clara, CA, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$168,000.0 - $258,750.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Systems Engineering Computer Clusters Data Centers General-Purpose Computing on Graphics Processing Units Monitoring of Systems Modbus Prometheus Software Requirements Analysis Supercomputing Systems Integration High Performance Computing
+4 more
Grafana Deep Learning Bacnet Splunk

Job description

NVIDIA is looking for a highly-motivated Technical Program Manager (TPM) to join our Applied Systems Engineering Team to drive datacenter integration for the next generation of NVIDIA AI supercomputing systems. This TPM will play a crucial role throughout the lifecycle of the latest AI systems at scale, from datacenter design and requirements definition, through systems integration of AI clusters into the datacenter environment, and support for these systems as they enter production.

This role will drive collaboration between engineering leaders across multiple hardware and software teams, helping us work together to build AI supercomputers for NVIDIA engineers and develop reference architectures to advise customers and partners.

What you’ll be doing:

  • Collaborate with outstanding engineers and architects to build and deploy large scale GPU computing systems based on NVIDIA’s reference supercomputing architectures

  • Lead the integration of new AI clusters with datacenter facilities with demanding requirements on power, cooling, and instrumentation

  • Coordinate design and fit-out of new datacenter builds, working with both internal engineering teams and external contractors

  • Own and produce detailed documentation for the end-to-end process for datacenter fit-out and integration

  • Communicate internally with engineering leadership to prioritize and address key issues essential to the success of our largest customers

Requirements

  • BS in Applied Science or Engineering (or equivalent experience)

  • 8+ years of overall experience

  • Experience with high-performance computing systems and GPU clusters deployed in on-premises datacenters

  • A passion for understanding challenging technical problems and driving the process of finding a solution

  • Strong teamwork and interpersonal skills, to facilitate building a collaborative workflow for coordination between many teams

Ways to stand out from the crowd:

  • Understanding of datacenter design, including familiarity with power and cooling technologies

  • Expertise in system monitoring and instrumentation of large clusters, using technologies such as Prometheus, Grafana, Splunk, Modbus, and BACNet

  • Experience working with the engineering or academic research community supporting high-performance computing or deep learning

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 258,750 USD for Level 4, and 200,000 USD - 322,000 USD for Level 5.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

4:19 min

Introduction to network security and endpoint monitoring architectures

Christoph Ruggenthaler · LIVE

2:08 min

History and scale of NVIDIA GPU computing

Paul Graham Paul Graham · LIVE

Videos

See all

Related articles

See all