Lead GPU Cluster Solutions Architect
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+8 more
Job description
- Design sophisticated GPU clusters from client requirements through production-ready architecture.
- Work directly with NVIDIA Reference Architecture, high-speed networking, storage, connectivity, and availability strategy.
- Solve challenging infrastructure problems where performance, reliability, power, cooling, and hardware constraints all matter.
- Have significant technical ownership over designs supporting enterprise and neocloud deployments.
- Collaborate closely with deployment, data center, supply chain, and program leadership to turn architecture into real-world infrastructure.
- Join a fast-growing AI infrastructure environment where your technical decisions directly impact customer outcomes.
- Competitive bonus and equity opportunity in addition to base compensation., * Own end-to-end technical architecture for GPU cluster deployments from customer requirements through deployment-ready design.
- Design GPU cluster configurations spanning compute, storage, networking, software requirements, and supporting infrastructure.
- Translate client technical requirements into complete bills of design covering all required compute, storage, networking, and connectivity components.
- Apply NVIDIA Reference Architecture principles, including HGX and NVL72-based GPU cluster designs.
- Design high-performance network fabrics using InfiniBand, RoCE, and high-speed Ethernet based on workload and performance requirements.
- Incorporate internet, VPN, firewall, dedicated circuit, protected optical, and other connectivity requirements into cluster architectures.
- Develop hot and cold sparing strategies designed to meet contracted availability and SLA commitments.
- Adapt cluster designs to site-specific power, cooling, space, hardware, and deployment constraints.
- Partner with data center teams to account for real-world facility limitations when finalizing technical architecture.
- Work with Supply Chain to ensure architecture decisions align with realistic hardware availability and lead times.
- Partner with deployment leadership and program management to translate designs into executable build plans.
- Support acceptance test planning and define technical criteria that validate the deployed architecture against the approved design.
- Evaluate and incorporate high-speed shared storage solutions such as Weka, VAST Data, and DDN where appropriate.
- Maintain technical ownership of architecture decisions while balancing performance, availability, cost, schedule, and operational supportability.
Requirements
Note: Must have 7+ years of directly relevant experience in solutions architecture, network engineering, or systems engineering supporting GPU, HPC, or large-scale compute infrastructure. Candidates must have hands-on GPU cluster design experience plus strong InfiniBand, RoCE, or high-speed Ethernet networking expertise. Generic cloud architecture or enterprise networking experience without meaningful GPU/HPC infrastructure exposure will not meet the requirements., * 7+ years of experience in solutions architecture, network engineering, systems engineering, or similar roles supporting GPU, HPC, or large-scale compute infrastructure.
- Deep working knowledge of NVIDIA Reference Architecture and GPU cluster design principles.
- Hands-on experience designing InfiniBand, RoCE, and/or high-speed Ethernet fabrics.
- Proven experience designing GPU or HPC clusters rather than solely consuming cloud infrastructure.
- Experience developing sparing and spares strategies for mission-critical infrastructure.
- Experience integrating firewalls, VPNs, dedicated circuits, protected optical connectivity, and related networking requirements into infrastructure designs.
- Experience with high-speed shared storage technologies such as Weka, VAST Data, or DDN.
- Strong understanding of compute, storage, networking, and data center infrastructure dependencies.
- Ability to translate complex customer requirements into complete, practical, buildable technical architectures.
- Strong cross-functional communication and documentation skills.
- Experience supporting enterprise customers or neocloud deployments is preferred.
- NVIDIA NCP program or certification experience is a plus.
- Experience with capacity planning or sparing modeling tools is a plus., Acceptance Testing, Artificial Intelligence (AI), Broadband, Capacity Management, Cloud Computing, Cross-Functional, Customer Support/Service, Design Services, Documentation, Ethernet, Firewalls, GPU (Graphics Processing Unit), Leadership, Network Architecture/Engineering, Network Connectivity, Network Operations Center, Network Software, Operational Support, Problem Solving Skills, Project/Program Management, Retirement Plan, Return on Capital Employed (ROCE), Service Level Agreement (SLA), Software Administration, Supply Chain, Systems Engineering, Technical/Engineering Design, VPN (Virtual Private Network), Willing to Travel, Work From Home
Benefits & conditions
Why You Will Love Working Here
- Work on technically challenging GPU infrastructure projects at the center of the AI compute market.
- Own architecture decisions that directly influence performance, reliability, scalability, and customer success.
- Gain exposure to cutting-edge NVIDIA GPU architectures and high-speed networking technologies.
- Collaborate with experienced infrastructure, deployment, data center, supply chain, and executive teams.
- Work remotely while remaining closely connected to real-world data center deployments.
- Opportunity to help establish repeatable architecture standards as the business scales.
- Competitive base compensation plus bonus and equity.
- Make a visible impact in a high-growth environment where strong technical judgment is valued., * Dental insurance
- Life insurance
- Paid time off
- Retirement plan
- Vision insurance
About the company
We are building next-generation AI infrastructure that gives enterprises access to high-performance GPU compute with speed, flexibility, and reliability. Our technical teams design and deploy sophisticated GPU clusters across data center environments, and we are looking for an architect who can turn demanding customer requirements into robust, buildable infrastructure. Confidential Employer.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
7 Cloud Computing Trends Coming in 2025 for Developers
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?