Architect AI Infrastructure & Fabric

VST Consulting, Inc
Plano, TX, United States
8 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Airflow Network Congestion Data Centers Linux Ethernet Firmware General Parallel File Systems InfiniBand Subnetting PCI Express Remote Direct Memory Access
+6 more
Weka AI Infrastructure Storage Technologies Bare Metal Hardware Infrastructure Nvme

Job description

We are building a GPU-as-a-Service and AI Factory practice supporting enterprise and industrial customers. This role focuses on the infrastructure beneath the operating system, including GPU node architecture, compute and storage fabrics, bare-metal provisioning, high-performance storage, and data center infrastructure.

You will be responsible for designing scalable GPU infrastructure and troubleshooting complex cluster, storage, networking, and performance issues., * Design GPU cluster physical and logical topology, including node configurations, rail-optimized fabric layouts, oversubscription ratios, and failure domains.

  • Architect and deploy InfiniBand NDR/XDR and high-performance Ethernet fabrics, including subnet management, UFM, adaptive routing, congestion control, and SHARP offload.
  • Design high-performance AI storage architectures using parallel filesystems such as WEKA, Lustre, GPFS, VAST, and DDN.
  • Work with GPUDirect Storage, NVMe-oF, and NFS-over-RDMA and size storage for dataloader reads, checkpoint writes, and artifact serving.
  • Own bare-metal cluster lifecycle, including provisioning, firmware and driver baselines, imaging, node validation, and burn-in.
  • Produce BOMs and infrastructure sizing for compute, networking, optics, and storage.
  • Validate infrastructure against customer power, cooling, floor loading, and deployment requirements.
  • Run scaling and fabric benchmarks including NCCL bus bandwidth, IB performance tests, IOR, and fio.
  • Diagnose interconnect and GPU cluster scaling issues, including link errors, topology binding, NUMA/PCIe affinity, GPUDirect RDMA, and I/O stalls.
  • Establish infrastructure standards for offshore delivery teams and review their work before customer delivery.

Requirements

Experience: 6+ Years HPC or AI Infrastructure Engineering Interview Mode: Virtual Practice: AI Infrastructure / GPU-as-a-Service, * 6+ years of HPC or AI infrastructure engineering experience.

  • Production multi-node GPU cluster experience is mandatory.
  • Deep experience with InfiniBand and/or high-performance Ethernet, including fabric design, subnet management, congestion behavior, and troubleshooting.
  • Strong experience with NVIDIA GPU platforms, including HGX or DGX-class systems, NVLink, NVSwitch, drivers, and firmware.
  • Experience designing and tuning parallel or scale-out filesystems for I/O-intensive workloads.
  • Strong bare-metal cluster provisioning and Linux systems engineering experience.
  • Understanding of data center infrastructure, including rack power, airflow, direct-liquid cooling concepts, cabling, and optics planning.

Nice to Have

  • NVIDIA Enterprise Reference Architecture experience.
  • Spectrum-X or BlueField DPU experience.
  • Liquid-cooled GPU deployment experience.
  • WEKA, VAST, or DDN certification.
  • NVIDIA networking certification.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:05 min

Acquiring Mellanox to build cohesive AI factories

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

8:22 min

Simulating a Linux terminal and running Spring Boot

Jakov Semenski · LIVE

Videos

See all

Related articles

See all