Architect AI Infrastructure & Fabric
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+6 more
Job description
We are building a GPU-as-a-Service and AI Factory practice supporting enterprise and industrial customers. This role focuses on the infrastructure beneath the operating system, including GPU node architecture, compute and storage fabrics, bare-metal provisioning, high-performance storage, and data center infrastructure.
You will be responsible for designing scalable GPU infrastructure and troubleshooting complex cluster, storage, networking, and performance issues., * Design GPU cluster physical and logical topology, including node configurations, rail-optimized fabric layouts, oversubscription ratios, and failure domains.
- Architect and deploy InfiniBand NDR/XDR and high-performance Ethernet fabrics, including subnet management, UFM, adaptive routing, congestion control, and SHARP offload.
- Design high-performance AI storage architectures using parallel filesystems such as WEKA, Lustre, GPFS, VAST, and DDN.
- Work with GPUDirect Storage, NVMe-oF, and NFS-over-RDMA and size storage for dataloader reads, checkpoint writes, and artifact serving.
- Own bare-metal cluster lifecycle, including provisioning, firmware and driver baselines, imaging, node validation, and burn-in.
- Produce BOMs and infrastructure sizing for compute, networking, optics, and storage.
- Validate infrastructure against customer power, cooling, floor loading, and deployment requirements.
- Run scaling and fabric benchmarks including NCCL bus bandwidth, IB performance tests, IOR, and fio.
- Diagnose interconnect and GPU cluster scaling issues, including link errors, topology binding, NUMA/PCIe affinity, GPUDirect RDMA, and I/O stalls.
- Establish infrastructure standards for offshore delivery teams and review their work before customer delivery.
Requirements
Experience: 6+ Years HPC or AI Infrastructure Engineering Interview Mode: Virtual Practice: AI Infrastructure / GPU-as-a-Service, * 6+ years of HPC or AI infrastructure engineering experience.
- Production multi-node GPU cluster experience is mandatory.
- Deep experience with InfiniBand and/or high-performance Ethernet, including fabric design, subnet management, congestion behavior, and troubleshooting.
- Strong experience with NVIDIA GPU platforms, including HGX or DGX-class systems, NVLink, NVSwitch, drivers, and firmware.
- Experience designing and tuning parallel or scale-out filesystems for I/O-intensive workloads.
- Strong bare-metal cluster provisioning and Linux systems engineering experience.
- Understanding of data center infrastructure, including rack power, airflow, direct-liquid cooling concepts, cabling, and optics planning.
Nice to Have
- NVIDIA Enterprise Reference Architecture experience.
- Spectrum-X or BlueField DPU experience.
- Liquid-cooled GPU deployment experience.
- WEKA, VAST, or DDN certification.
- NVIDIA networking certification.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Stephan Gillich - Bringing AI Everywhere
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
7 Cloud Computing Trends Coming in 2025 for Developers
Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?