> Markdown version of [/jobs/ext/2722699-ai-cluster-architect](https://www.wearedevelopers.com/jobs/ext/2722699-ai-cluster-architect). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Cluster Architect - **Company:** The Constant Company, LLC - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $165,000.0 - $185,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Common Lisp Object Systems, Computer Clusters, Digital Architecture, Microprocessors, Network Interface Controllers, InfiniBand, Network Architecture, PCI Express, Graphics Processing Unit (GPU), 3-tier Architectures, Network Server, Server Operating Systems & Platforms - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/ai-cluster-architect-vultr-com-8958422 ## About the Role * 7+ years designing or building large-scale HPC, AI, or hyperscale GPU clusters. * Expert understanding of GPU and accelerator system design, including node topology, PCIe/NVLink/NVSwitch/ROCm, and NIC-to-GPU affinity considerations. * Strong familiarity with InfiniBand, RoCE, and SpectrumX networking, including multi-tier, multi-plane, Clos/dragonfly variants, and large-radix switch design. * Demonstrated experience modeling power draw and thermal characteristics of servers, GPUs, NICs, switches, optics, and storage systems. * Ability to design networks that maintain full non-blocking performance or intentionally introduce over/under-subscription while understanding impacts on workload performance. * Proven ability to gather and analyze vendor SKU-level specifications and incorporate them into scalable cluster architectures. * Experience balancing customer-driven requirements for compute, storage, and service density in combination with overall GPU count. * Strong documentation, communication, and cross-functional collaboration skills., This salary can vary based on location, years of experience, background and skill set. ## Description The architect must understand how different GPU SKUs, NICs, switches, and fabrics interact at scale, including their individual and aggregate power and thermal characteristics. They will evaluate multi-plane, rail-optimized, and tiered fabric designs across technologies like InfiniBand, RoCE, and SpectrumX to ensure the networking architecture supports the intended GPU count without overrunning facility limits or switch radix and/or topology constraints. This role balances customer-specific requirements for compute, storage, and service density, ensuring that the final cluster design maintains acceptable levels of GPU and fabric performance, while maximizing the number of usable GPUs within the total power budget., * Architect large-scale GPU clusters within fixed site power budgets that optimizes for maximum GPU density while reserving necessary headroom for compute services, storage, and networking. * Model and validate power consumption across the full cluster bill of materials (GPUs, CPUs, NICs, switches, fabric components, storage, and facility limits). * Evaluate tradeoffs across multiple fabric networking architectures (InfiniBand, RoCE, SpectrumX) as well as multi-plane, 2-tier/3-tier, and rail-optimized topologies. * Determine network scale limits based on switch radix, link speed, topology, and blocking requirements. * Gather, interpret, and maintain detailed SKU-level power and thermal specifications for GPUs, NICs, switches, DPUs, storage, and server platforms. * Develop power-aware cluster configuration templates and capacity-planning models that can scale across sites with varying constraints and allow for quick iteration and ideation. * Document architecture, design choices, tradeoff analyses, and operational considerations for deployment and lifecycle management. * Provide guidance on future-proofing, including the ability to incorporate next-gen GPUs, NICs, or fabrics. * Collaborate with vendors on novel fabric architectures that enable large-scale cluster deployments (100k+ GPUs) ## Related Videos - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Kubernetes Security - Challenge and Opportunity](https://www.wearedevelopers.com/videos/412-kubernetes-security-challenge-and-opportunity) - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) - [tRPC: API schemas are pure overhead](https://www.wearedevelopers.com/videos/796-trpc-api-schemas-are-pure-overhead) - [Building Better with Nuxt 3](https://www.wearedevelopers.com/videos/320-building-better-with-nuxt-3) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)