Solutions Architect, Infrastructure

NVIDIA Ltd.
Redmond, WA, United States
4 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$152,000.0 - $241,500.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Application Layers Intelligent Platform Management Interface Big Data BIOS Cloud Computing Cloud Engineering Nvidia CUDA Computer Engineering Network Congestion Data Centers Software Debugging
+18 more
Linux Distributed Computing Environment Distributed Systems Firmware PCI Express Performance Tuning Remote Direct Memory Access Cloud Services Remote Infrastructure Management Subsystems System Software Graphics Processing Unit (GPU) Computer Networking Systems Extensible Firmware Interface Large Language Models Deep Learning Information Technology Hardware Infrastructure

Job description

Do you thrive on taking a strategic product from launch to go-to-market at scale across the world’s largest customers? NVIDIA is looking for an Infrastructure Solutions Architect to lead deployment and bring-up of our next-generation Data Center GPUs and networking platforms.

As part of the NVIDIA Solutions Architecture team, we navigate uncharted technical and organizational spaces - serving as the bridge between early platform readiness, cloud engineering teams, product strategy, and large-scale customer deployments. We are looking for Solution Architects to combine hands-on infrastructure expertise with multi-functional leadership to accelerate adoption of NVIDIA technologies across worldwide cloud hosting providers and large enterprise environments.

What You’ll Be Doing:

  • Lead end-to-end execution for Hyperscaler customers to rapidly bring NVIDIA Data Center GPU and networking platforms to market at scale.

  • Drive strategic partnership and alignment with Product teams to understand roadmap intent, co-define critical metrics, and ensure unified direction across technical, sales, and leadership organizations.

  • Influence without authority across Product, Engineering, Sales, Operations, and CSP customers, driving clarity, alignment, and unblock paths for scale-up.

  • Analyze deployment and performance data, identifying product health trends, system bottlenecks, and operational risks.

  • Solve challenging technical problems involving GPUs, networking, drivers, containers, firmware, and distributed system interactions.

  • Deliver streamlined executive-level communication on status, risks, progress, and required decisions.

  • Collaborate with Product and Engineering, enabling future improvements in platform design, validation, and operational workflows.

Requirements

  • BS/MS/PhD in Electrical/Computer Engineering, Computer Science, Physics, or similar, or equivalent experience.

  • 4+ years experience in Solutions Architecture, Infrastructure Engineering, or similar technical roles.

  • Hands-on experience with bring-up and validation of large-scale NVIDIA GPU platforms, including multi-GPU and multi-node architectures.

  • Understanding of high-performance networking technologies (e.g., RDMA, congestion control, high-bandwidth interconnects) and their role in distributed AI workloads.

  • Familiarity with NVIDIA system software stacks: CUDA, NCCL, NVSwitch/NVLink, driver behavior, and performance tuning.

  • Proficiency with Linux systems tools for identifying issues and evaluating system performance, such as: dmesg, journalctl, lspci, numactl, ethtool, iostat, perf, nvidia-smi, top/htop, ipmitool, container-level tooling, and related utilities.

  • Understanding of server hardware architecture, including PCIe topologies, system firmware, NUMA, BIOS/UEFI configuration, power/thermal envelopes, and memory/subsystem behavior.

  • Understanding of BMC/IPMI/Redfish for remote management, hardware health monitoring, and out-of-band debugging during early-stage bring-up.

  • Strong Linux fundamentals across drivers, kernel subsystems, cgroups, containers, and node-level performance analysis.

  • Ability to identify performance bottlenecks at the cluster, node, accelerator, network, or application layer.

Ways to Stand Out from the Crowd:

  • Outstanding interpersonal skills and the ability to build clarity and direction across diverse, fast paced technical teams.

  • Knowledge of Compute and networking infrastructure (e.g., Instance types, networking primitives, high-performance communication paths etc) at Hyperscalers or Cloud Service Providers.

  • Demonstrated leadership resolving multi-team infrastructure challenges across engineering, product, and customer groups.

  • A consistent record of taking GPU or infrastructure products from pilot to high-volume deployment in large data center environments.

  • Familiarity with modern deep learning, LLM architectures, and distributed training/inference challenges at scale.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

41 sec

Massive client data loss and bio-digital storage

Chris Heilmann +1 · LIVE

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

Videos

See all

Related articles

See all