Lead GPU Cluster Solutions Architect

NVIDIA Ltd.
Pittsburgh, PA, United States
14 days ago
Apply on www.careerbuilder.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Compensation
$140,000.0 - $240,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Systems Engineering Cloud Computing Cloud Engineering Computer Clusters Computer Networks Data Centers Ethernet InfiniBand Virtual Private Networks (VPN) Network Architecture Network Connections
+8 more
Server Administration Software Requirements Analysis Weka AI Infrastructure Graphics Processing Unit (GPU) Computer Network Operations Firewalls (Computer Science) Storage Technologies

Job description

  • Design sophisticated GPU clusters from client requirements through production-ready architecture.
  • Work directly with NVIDIA Reference Architecture, high-speed networking, storage, connectivity, and availability strategy.
  • Solve challenging infrastructure problems where performance, reliability, power, cooling, and hardware constraints all matter.
  • Have significant technical ownership over designs supporting enterprise and neocloud deployments.
  • Collaborate closely with deployment, data center, supply chain, and program leadership to turn architecture into real-world infrastructure.
  • Join a fast-growing AI infrastructure environment where your technical decisions directly impact customer outcomes.
  • Competitive bonus and equity opportunity in addition to base compensation., * Own end-to-end technical architecture for GPU cluster deployments from customer requirements through deployment-ready design.
  • Design GPU cluster configurations spanning compute, storage, networking, software requirements, and supporting infrastructure.
  • Translate client technical requirements into complete bills of design covering all required compute, storage, networking, and connectivity components.
  • Apply NVIDIA Reference Architecture principles, including HGX and NVL72-based GPU cluster designs.
  • Design high-performance network fabrics using InfiniBand, RoCE, and high-speed Ethernet based on workload and performance requirements.
  • Incorporate internet, VPN, firewall, dedicated circuit, protected optical, and other connectivity requirements into cluster architectures.
  • Develop hot and cold sparing strategies designed to meet contracted availability and SLA commitments.
  • Adapt cluster designs to site-specific power, cooling, space, hardware, and deployment constraints.
  • Partner with data center teams to account for real-world facility limitations when finalizing technical architecture.
  • Work with Supply Chain to ensure architecture decisions align with realistic hardware availability and lead times.
  • Partner with deployment leadership and program management to translate designs into executable build plans.
  • Support acceptance test planning and define technical criteria that validate the deployed architecture against the approved design.
  • Evaluate and incorporate high-speed shared storage solutions such as Weka, VAST Data, and DDN where appropriate.
  • Maintain technical ownership of architecture decisions while balancing performance, availability, cost, schedule, and operational supportability.

Requirements

Note: Must have 7+ years of directly relevant experience in solutions architecture, network engineering, or systems engineering supporting GPU, HPC, or large-scale compute infrastructure. Candidates must have hands-on GPU cluster design experience plus strong InfiniBand, RoCE, or high-speed Ethernet networking expertise. Generic cloud architecture or enterprise networking experience without meaningful GPU/HPC infrastructure exposure will not meet the requirements., * 7+ years of experience in solutions architecture, network engineering, systems engineering, or similar roles supporting GPU, HPC, or large-scale compute infrastructure.

  • Deep working knowledge of NVIDIA Reference Architecture and GPU cluster design principles.
  • Hands-on experience designing InfiniBand, RoCE, and/or high-speed Ethernet fabrics.
  • Proven experience designing GPU or HPC clusters rather than solely consuming cloud infrastructure.
  • Experience developing sparing and spares strategies for mission-critical infrastructure.
  • Experience integrating firewalls, VPNs, dedicated circuits, protected optical connectivity, and related networking requirements into infrastructure designs.
  • Experience with high-speed shared storage technologies such as Weka, VAST Data, or DDN.
  • Strong understanding of compute, storage, networking, and data center infrastructure dependencies.
  • Ability to translate complex customer requirements into complete, practical, buildable technical architectures.
  • Strong cross-functional communication and documentation skills.
  • Experience supporting enterprise customers or neocloud deployments is preferred.
  • NVIDIA NCP program or certification experience is a plus.
  • Experience with capacity planning or sparing modeling tools is a plus., Acceptance Testing, Artificial Intelligence (AI), Broadband, Capacity Management, Cloud Computing, Cross-Functional, Customer Support/Service, Design Services, Documentation, Ethernet, Firewalls, GPU (Graphics Processing Unit), Leadership, Network Architecture/Engineering, Network Connectivity, Network Operations Center, Network Software, Operational Support, Problem Solving Skills, Project/Program Management, Retirement Plan, Return on Capital Employed (ROCE), Service Level Agreement (SLA), Software Administration, Supply Chain, Systems Engineering, Technical/Engineering Design, VPN (Virtual Private Network), Willing to Travel, Work From Home

Benefits & conditions

Why You Will Love Working Here

  • Work on technically challenging GPU infrastructure projects at the center of the AI compute market.
  • Own architecture decisions that directly influence performance, reliability, scalability, and customer success.
  • Gain exposure to cutting-edge NVIDIA GPU architectures and high-speed networking technologies.
  • Collaborate with experienced infrastructure, deployment, data center, supply chain, and executive teams.
  • Work remotely while remaining closely connected to real-world data center deployments.
  • Opportunity to help establish repeatable architecture standards as the business scales.
  • Competitive base compensation plus bonus and equity.
  • Make a visible impact in a high-growth environment where strong technical judgment is valued., * Dental insurance
  • Life insurance
  • Paid time off
  • Retirement plan
  • Vision insurance

About the company

We are building next-generation AI infrastructure that gives enterprises access to high-performance GPU compute with speed, flexibility, and reliability. Our technical teams design and deploy sophisticated GPU clusters across data center environments, and we are looking for an architect who can turn demanding customer requirements into robust, buildable infrastructure. Confidential Employer.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:59 min

Scaling up and scaling out GPU clusters

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

1:57 min

Routing cross-rack traffic seamlessly with NCCL

Kevin Klues Kevin Klues · World Congress 2025

4:52 min

Connecting namespaces with local virtual ethernet pairs

Oliver Seitz Oliver Seitz · World Congress 2025

1:51 min

Bypassing the CPU stack with remote direct memory access

Lerna Ekmekcioglu Lerna Ekmekcioglu · Europe 2026 Virtual

3:23 min

The AI workload technology stack and its components

Lerna Ekmekcioglu Lerna Ekmekcioglu · Europe 2026 Virtual

2:12 min

Implementing automotive ethernet and connected remote vehicle telemetry applications

David Romić · World Congress 2023

Videos

See all

Related articles

See all