> Markdown version of [/jobs/ext/2167029-lead-gpu-cluster-solutions-architect](https://www.wearedevelopers.com/jobs/ext/2167029-lead-gpu-cluster-solutions-architect). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Lead GPU Cluster Solutions Architect - **Company:** NVIDIA Ltd. - **Location:** Pittsburgh, PA, United States (Remote available) - **Experience:** Expert - **Salary:** $140,000.0 - $240,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Cloud Computing, Cloud Engineering, Computer Clusters, Computer Networks, Data Centers, Ethernet, InfiniBand, Virtual Private Networks (VPN), Network Architecture, Network Connections, Server Administration, Software Requirements Analysis, Weka, AI Infrastructure, Graphics Processing Unit (GPU), Computer Network Operations, Firewalls (Computer Science), Storage Technologies - **Published:** August 21, 2026 - **Apply:** https://www.careerbuilder.com/job-details/lead-gpu-cluster-solutions-architect-pittsburgh-pa--254208f6-391b-4210-82b1-755b8ee3f9c5 ## About the Role Note: Must have 7+ years of directly relevant experience in solutions architecture, network engineering, or systems engineering supporting GPU, HPC, or large-scale compute infrastructure. Candidates must have hands-on GPU cluster design experience plus strong InfiniBand, RoCE, or high-speed Ethernet networking expertise. Generic cloud architecture or enterprise networking experience without meaningful GPU/HPC infrastructure exposure will not meet the requirements., * 7+ years of experience in solutions architecture, network engineering, systems engineering, or similar roles supporting GPU, HPC, or large-scale compute infrastructure. * Deep working knowledge of NVIDIA Reference Architecture and GPU cluster design principles. * Hands-on experience designing InfiniBand, RoCE, and/or high-speed Ethernet fabrics. * Proven experience designing GPU or HPC clusters rather than solely consuming cloud infrastructure. * Experience developing sparing and spares strategies for mission-critical infrastructure. * Experience integrating firewalls, VPNs, dedicated circuits, protected optical connectivity, and related networking requirements into infrastructure designs. * Experience with high-speed shared storage technologies such as Weka, VAST Data, or DDN. * Strong understanding of compute, storage, networking, and data center infrastructure dependencies. * Ability to translate complex customer requirements into complete, practical, buildable technical architectures. * Strong cross-functional communication and documentation skills. * Experience supporting enterprise customers or neocloud deployments is preferred. * NVIDIA NCP program or certification experience is a plus. * Experience with capacity planning or sparing modeling tools is a plus., Acceptance Testing, Artificial Intelligence (AI), Broadband, Capacity Management, Cloud Computing, Cross-Functional, Customer Support/Service, Design Services, Documentation, Ethernet, Firewalls, GPU (Graphics Processing Unit), Leadership, Network Architecture/Engineering, Network Connectivity, Network Operations Center, Network Software, Operational Support, Problem Solving Skills, Project/Program Management, Retirement Plan, Return on Capital Employed (ROCE), Service Level Agreement (SLA), Software Administration, Supply Chain, Systems Engineering, Technical/Engineering Design, VPN (Virtual Private Network), Willing to Travel, Work From Home ## Description * Design sophisticated GPU clusters from client requirements through production-ready architecture. * Work directly with NVIDIA Reference Architecture, high-speed networking, storage, connectivity, and availability strategy. * Solve challenging infrastructure problems where performance, reliability, power, cooling, and hardware constraints all matter. * Have significant technical ownership over designs supporting enterprise and neocloud deployments. * Collaborate closely with deployment, data center, supply chain, and program leadership to turn architecture into real-world infrastructure. * Join a fast-growing AI infrastructure environment where your technical decisions directly impact customer outcomes. * Competitive bonus and equity opportunity in addition to base compensation., * Own end-to-end technical architecture for GPU cluster deployments from customer requirements through deployment-ready design. * Design GPU cluster configurations spanning compute, storage, networking, software requirements, and supporting infrastructure. * Translate client technical requirements into complete bills of design covering all required compute, storage, networking, and connectivity components. * Apply NVIDIA Reference Architecture principles, including HGX and NVL72-based GPU cluster designs. * Design high-performance network fabrics using InfiniBand, RoCE, and high-speed Ethernet based on workload and performance requirements. * Incorporate internet, VPN, firewall, dedicated circuit, protected optical, and other connectivity requirements into cluster architectures. * Develop hot and cold sparing strategies designed to meet contracted availability and SLA commitments. * Adapt cluster designs to site-specific power, cooling, space, hardware, and deployment constraints. * Partner with data center teams to account for real-world facility limitations when finalizing technical architecture. * Work with Supply Chain to ensure architecture decisions align with realistic hardware availability and lead times. * Partner with deployment leadership and program management to translate designs into executable build plans. * Support acceptance test planning and define technical criteria that validate the deployed architecture against the approved design. * Evaluate and incorporate high-speed shared storage solutions such as Weka, VAST Data, and DDN where appropriate. * Maintain technical ownership of architecture decisions while balancing performance, availability, cost, schedule, and operational supportability. ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [A Deep Dive on How To Leverage the NVIDIA GB200 for Ultra-Fast Training and Inference on Kubernetes](https://www.wearedevelopers.com/videos/1625-a-deep-dive-on-how-to-leverage-the-nvidia-gb200-for-ultra-fast-training-and-inference-on-kubernetes) - [Building a hypercar from scratch](https://www.wearedevelopers.com/videos/607-building-a-hypercar-from-scratch) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)