> Markdown version of [/jobs/ext/2976592-architect-ai-infrastructure-fabric](https://www.wearedevelopers.com/jobs/ext/2976592-architect-ai-infrastructure-fabric). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Architect AI Infrastructure & Fabric - **Company:** VST Consulting, Inc - **Location:** Plano, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Network Congestion, Data Centers, Linux, Ethernet, Firmware, General Parallel File Systems, InfiniBand, Subnetting, PCI Express, Remote Direct Memory Access, Weka, AI Infrastructure, Storage Technologies, Bare Metal, Hardware Infrastructure, Nvme - **Published:** September 18, 2026 - **Apply:** https://www.dice.com/job-detail/8e1d2f9d-3686-44c8-95a0-c07d769af26b ## About the Role Experience: 6+ Years HPC or AI Infrastructure Engineering Interview Mode: Virtual Practice: AI Infrastructure / GPU-as-a-Service, * 6+ years of HPC or AI infrastructure engineering experience. * Production multi-node GPU cluster experience is mandatory. * Deep experience with InfiniBand and/or high-performance Ethernet, including fabric design, subnet management, congestion behavior, and troubleshooting. * Strong experience with NVIDIA GPU platforms, including HGX or DGX-class systems, NVLink, NVSwitch, drivers, and firmware. * Experience designing and tuning parallel or scale-out filesystems for I/O-intensive workloads. * Strong bare-metal cluster provisioning and Linux systems engineering experience. * Understanding of data center infrastructure, including rack power, airflow, direct-liquid cooling concepts, cabling, and optics planning. Nice to Have * NVIDIA Enterprise Reference Architecture experience. * Spectrum-X or BlueField DPU experience. * Liquid-cooled GPU deployment experience. * WEKA, VAST, or DDN certification. * NVIDIA networking certification. ## Description We are building a GPU-as-a-Service and AI Factory practice supporting enterprise and industrial customers. This role focuses on the infrastructure beneath the operating system, including GPU node architecture, compute and storage fabrics, bare-metal provisioning, high-performance storage, and data center infrastructure. You will be responsible for designing scalable GPU infrastructure and troubleshooting complex cluster, storage, networking, and performance issues., * Design GPU cluster physical and logical topology, including node configurations, rail-optimized fabric layouts, oversubscription ratios, and failure domains. * Architect and deploy InfiniBand NDR/XDR and high-performance Ethernet fabrics, including subnet management, UFM, adaptive routing, congestion control, and SHARP offload. * Design high-performance AI storage architectures using parallel filesystems such as WEKA, Lustre, GPFS, VAST, and DDN. * Work with GPUDirect Storage, NVMe-oF, and NFS-over-RDMA and size storage for dataloader reads, checkpoint writes, and artifact serving. * Own bare-metal cluster lifecycle, including provisioning, firmware and driver baselines, imaging, node validation, and burn-in. * Produce BOMs and infrastructure sizing for compute, networking, optics, and storage. * Validate infrastructure against customer power, cooling, floor loading, and deployment requirements. * Run scaling and fabric benchmarks including NCCL bus bandwidth, IB performance tests, IOR, and fio. * Diagnose interconnect and GPU cluster scaling issues, including link errors, topology binding, NUMA/PCIe affinity, GPUDirect RDMA, and I/O stalls. * Establish infrastructure standards for offshore delivery teams and review their work before customer delivery. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [AI Factories at Scale](https://www.wearedevelopers.com/videos/1139-ai-factories-at-scale) - [Discover the open source trio you didn’t expect: .NET and PostgreSQL on Linux](https://www.wearedevelopers.com/videos/2042-discover-the-open-source-trio-you-didn-t-expect-net-and-postgresql-on-linux) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)