AI Data Center Architect

NVIDIA Ltd.
Plano, TX, United States
1 day ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Computing Platforms Microsoft Azure Cloud Computing Computer Clusters Data Centers Distributed Computing Environment InfiniBand Machine Learning Performance Tuning AI Infrastructure
+8 more
Graphics Processing Unit (GPU) Google Cloud Large Language Models AI Platforms Kubernetes Storage Technologies Machine Learning Operations Hardware Infrastructure

Job description

We are seeking a highly skilled AI Data Center Architect with hands-on experience designing and architecting modern AI infrastructure at enterprise or hyperscale scale. This is not a traditional enterprise data center architecture role. The ideal candidate will have deep expertise in GPU infrastructure, AI Factories, GPU-as-a-Service (GPUaaS), NVIDIA reference architectures, high-performance networking, AI storage, Kubernetes, and large-scale AI/LLM workloads. The successful candidate will be responsible for designing end-to-end AI infrastructure solutions, including GPU compute, high-performance networking, storage, orchestration, multi-tenancy, and AI platform services. The individual should be comfortable working directly with customers and technical stakeholders to translate AI workload requirements into scalable infrastructure architectures., * Design and architect enterprise and hyperscale AI Factory environments from the ground up.

  • Develop scalable GPU infrastructure and GPU-as-a-Service (GPUaaS) architectures.
  • Design infrastructure for large-scale AI training, inference, LLM, and GenAI workloads.
  • Develop solutions based on NVIDIA Enterprise Reference Architectures.
  • Architect NVIDIA GPU platforms including DGX, HGX, H200, GB200, and Blackwell systems.
  • Design high-performance AI networking using InfiniBand, RoCE, Spectrum-X, NVLink, and BlueField DPUs.
  • Perform GPU cluster sizing, capacity planning, performance optimization, and scalability analysis.
  • Design high-performance storage architectures capable of supporting large-scale GPU and AI workloads.
  • Architect Kubernetes-based AI platforms, including GPU scheduling, workload orchestration, resource management, and multi-tenancy.
  • Design distributed training and inference infrastructure for modern LLM and GenAI workloads.
  • Develop architectures for AI Cloud, GPU Cloud, and on-premises AI Factory deployments.
  • Evaluate infrastructure requirements across compute, networking, storage, orchestration, and AI platform layers.
  • Work with engineering, infrastructure, cloud, networking, and AI/ML teams to develop end-to-end solutions.
  • Participate in customer-facing architecture discussions, technical workshops, solution design, and presentations.
  • Create high-level and low-level architecture designs, technical documentation, reference architectures, and solution proposals.
  • Evaluate emerging NVIDIA technologies and AI infrastructure trends and incorporate them into future architectures.

Requirements

  • 4+ years of experience designing AI-focused infrastructure, accelerated computing platforms, or GPU-based data center environments.
  • Proven experience architecting AI Factories, GPU clusters, or GPUaaS platforms.
  • Experience designing environments supporting 100+ GPUs is strongly preferred.
  • Deep understanding of NVIDIA AI infrastructure and Enterprise Reference Architectures.

Hands-on architecture experience with one or more of:

  • NVIDIA DGX
  • NVIDIA HGX
  • H200
  • GB200
  • NVIDIA Blackwell platforms

Strong understanding of:

  • InfiniBand
  • Spectrum-X
  • RoCE
  • NVLink / NVSwitch
  • BlueField DPUs
  • 400G/800G networking
  • Experience with GPU cluster sizing, performance optimization, scaling, and capacity planning.
  • Strong understanding of LLM training and inference infrastructure.
  • Experience with Kubernetes for AI/GPU workloads.
  • Knowledge of GPU scheduling, workload orchestration, and multi-tenant GPU environments.
  • Experience designing distributed AI/ML training infrastructure.
  • Strong knowledge of high-performance storage architectures for AI workloads.
  • Understanding of MLOps, AI platforms, and GenAI infrastructure.
  • Ability to design complete AI infrastructure solutions spanning compute, networking, storage, orchestration, and platform services., * Experience designing on-premises AI Factories, Sovereign AI infrastructure, or GPU Cloud platforms.
  • Experience with Azure, AWS, or Google Cloud AI infrastructure.
  • Experience with NVIDIA AI Enterprise and the broader NVIDIA AI software ecosystem.
  • Experience with large-scale distributed training frameworks and AI workload orchestration.
  • Experience with AI infrastructure benchmarking and performance optimization.
  • Customer-facing architecture, consulting, solution engineering, or technical pre-sales experience.
  • Experience developing technical proposals, architecture diagrams, bills of materials, and solution designs.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

51 sec

Repurposing hardware and operating underwater data centers

Chris Heilmann +1 · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

1:57 min

Routing cross-rack traffic seamlessly with NCCL

Kevin Klues Kevin Klues · World Congress 2025

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:03 min

Managing massive power consumption scaling in AI data centers

Stephan Gillich Stephan Gillich +3 · World Congress 2024

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · World Congress 2022

Videos

See all

Related articles

See all