Senior Solutions Architect, Cluster Design and Architecture - Networking

NVIDIA Corporation
Santa Clara, CA, United States
1 day ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$184,000.0 - $287,500.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Computing Platforms Computer Clusters Computer Engineering Network Congestion Distributed Computing Environment Distributed Systems Network Topologies InfiniBand Network Architecture Network Planning and Design Software Systems
+5 more
Supercomputing AI Infrastructure Graphics Processing Unit (GPU) Computer Network Technologies Information Technology

Job description

NVIDIA is building the world’s most groundbreaking and innovative accelerated computing platforms for AI and HPC. Because of our work, scientists, researchers, and engineers can push the boundaries of what’s possible. We pioneered a supercharged form of computing that powers everything from breakthrough AI research to the world’s fastest supercomputers. We are seeing a highly motivated Senior Solutions Architect to join the Cluster Design and Architecture team with a focus on networking technologies. As AI workloads scale to unprecedented levels, the network is the backbone that makes large compute clusters possible. In this role, you will be at the forefront of assisting with designs and architectures for next-generation networking solutions that connect thousands of GPUs and enable the world’s most advanced AI supercomputers and enterprise AI infrastructure in the field.

As a Solutions Architect, you will act as a key technical expert connecting NVIDIA’s new networking technology builds. These include Infiniband, Spectrum-X, NVLink, and all software solutions. You will work directly between engineering and field teams to support customers with fast paced requirements. You will work on end-to-end cluster design, network topology and architecture optimization, performance modeling and validation. Your expertise will directly influence how the world’s leading AI companies, cloud providers, hyperscalers, research institutions, and enterprises build their infrastructure.

What you’ll be doing:

  • Partner with internal engineering efforts in GPU cluster building and networking and convey architecture and guidelines information both direct to customer and with field teams supporting customers
  • Guide field teams and their customers in cluster design, weighing design principles but also complex, situational limitations to make the most performant and supportable GPU clusters possible
  • Work closely with field teams supporting customers to ensure successful first deployments with new products, including new network architectures and topologies
  • Feedback customer/field perspectives on networking development and workflows back to engineering teams building internal clusters and/or composing customer facing documentation on guidelines and service flows
  • Perform hands-on work to assist field teams debugging issues relating to network build, configuration, and performance, bringing to bear internal engineering expertise and known bugs

Requirements

  • BS, MS, or PhD in Computer Science, Electrical Engineering, Computer Engineering, Physics, or related field (or equivalent experience)
  • 8+ years of experience in network architecture, network design, network validation and troubleshooting
  • Proven expertise in designing large-scale distributed systems, AI clusters, or HPC infrastructure
  • Ability to translate complex engineering concepts into customer-ready documentation, diagrams, and reference material

Ways to stand out from the crowd:

  • Experience leading large-scale AI Factory or HPC cluster bring-ups or builds
  • Hands-on experience with NVIDIA networking products including, but not limited to, Infiniband, Spectrum-X, BlueField, etc.
  • Knowledge of NCCL, MPI, and collective communication patterns in distributed training as it pertains to networking patterns and design
  • Background in network performance optimization, congestion control, and validation at scale
  • External customer facing skill-set and background

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

About the company

NVIDIA is widely considered to be one of the technology world’s most desirable employers with very competitive benefits. We have some of the most forward-thinking and innovative people in the world working for us. If you’re creative and autonomous, we want to hear from you!

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:59 min

Scaling up and scaling out GPU clusters

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

2:04 min

Insights on transitioning from supercomputing to technical education

Andrew Holway · LIVE

2:36 min

Experiencing Conway's Law in heterogeneous software systems

Michael Jaeger Michael Jaeger · Europe 2026 Virtual

1:57 min

Routing cross-rack traffic seamlessly with NCCL

Kevin Klues Kevin Klues · World Congress 2025

1:44 min

Background and career journey in regulated software systems

Martin Hynie · Coffee With Developers

3:23 min

The AI workload technology stack and its components

Lerna Ekmekcioglu Lerna Ekmekcioglu · Europe 2026 Virtual

Videos

See all

Related articles

See all