Senior Data Center Connectivity Engineer

NVIDIA Ltd.
United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$208,000.0 - $333,500.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Cloud Computing Computer Clusters Data Centers InfiniBand Network Security Network Architecture Network Diagrams Remote Direct Memory Access High Performance Computing

Job description

The connectivity engineer translates product reference architectures and logical network diagrams into physical builds. This applies to NVIDIA’s AI Factory build guidelines and NVIDIA’s large-scale internal research clusters. This role will act as the lead engineer for all in-cluster cabling, pathway and rack layout optimizations required to power global-scale AI deployments, ensuring the cluster is co-designed with facilities infrastructure (Power&Cooling) and Infrastructure Software. This role provides an outstanding opportunity to be at the forefront of NVIDIA’s technology roadmap!

What you’ll be doing:

  • Own the development of connectivity reference designs based on requirements from cluster architecture, network engineering, infrastructure software and product hardware teams.

  • Build and develop comprehensive documentation, including detailed rack elevations and network architecture diagrams and cabling point-to-point list. Support projects throughout design and deployment phases.

  • Serve as the primary engineering support, closely collaborating with deployment and field teams to ensure successful cluster build-out and operation.

  • Strategically co-design the cluster with power and cooling infrastructure teams, ensuring a thorough understanding of all facility architectural requirements (Arch, power, cooling).

  • Work with hardware, network and security teams to translate software stack requirements into physical requirements: hardware selection, fault domain, network architecture.

  • Develop new solutions and products in the connectivity space to accelerate the deployment of large scale AI Factories

Requirements

  • Minimum of 12+ years in a connectivity, network architecture or engineering role within a Hyperscale Cloud Provider, large-scale enterprise data center, or High-Performance Computing (HPC) environment.

  • BA or BS (or equivalent experience).

  • Consistent record of designing, deploying, and operating network fabrics for thousands of GPU/CPU nodes.

  • Deep expertise in high-speed interconnect technologies, including InfiniBand, RoCE, and RDMA.

  • Proven experience designing connectivity solutions for high-density GPU clusters (100kW+ per rack) and understanding the unique front-end and back-end requirements for AI training vs. inference.

  • Deep understanding of data center infrastructure, including rack power/cooling, cable management, and physical density constraints.

  • Demonstrated ability to lead multidisciplinary teams and complete sophisticated technical initiatives.

Ways to stand out from the crowd:

  • Deep expertise with NVIDIA’s compute and network product families and deployment standards.

  • Comfortable operating at the intersection of network engineering, MEP systems, and Infrastructure as a Service software layer.

  • Experienced with field deployments and/or global reference design documentation, ideally both.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 208,000 USD - 333,500 USD.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:27 min

Introduction to WebAssembly in a cloud computing context

Edo Edo · WWC 2024

51 sec

Repurposing hardware and operating underwater data centers

Chris Heilmann +1 · LIVE

1:57 min

Routing cross-rack traffic seamlessly with NCCL

Kevin Klues Kevin Klues · WWC 2025

3:05 min

Acquiring Mellanox to build cohesive AI factories

Michael Kagan Michael Kagan +1 · WWC Europe 2026

2:51 min

Alibaba Cloud developer resources and cloud computing training

Cheng Zhang · LIVE

4:03 min

Managing massive power consumption scaling in AI data centers

Stephan Gillich Stephan Gillich +3 · WWC 2024

Videos

See all

Related articles

See all