GPU Solutions Architect - Cloud

Digital Dhara LLC
Santa Clara, CA, United States
18 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience required
1 year minimum
Working hours
Shift work
Job source

Tech stack

Bash Shell Cloud Engineering Computer Engineering Distributed Systems Ethernet Firmware General Parallel File Systems IBM Storage InfiniBand Python (Programming Language) Linux System Administration Octopus Deploy
+15 more
Reliability Engineering Ansible Prometheus Weka AI Infrastructure Scripting Cloud Platform System High Performance Computing Grafana Kubernetes Information Technology Data Management Slurm Hardware Infrastructure Terraform

Requirements

  • BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Mathematics, Physics, or related discipline (or equivalent experience).
  • 8+ years of experience in:

  • Production Infrastructure
  • Cloud Engineering
  • Solutions Architecture
  • Site Reliability Engineering (SRE)
  • HPC Environments
  • Similar technical disciplines OR

5+ years of exceptional specialist-level experience supporting large-scale GPU or AI infrastructure.

Technical Expertise: Experience building, operating, and optimizing distributed infrastructure in production environments. Deep hands-on expertise in one or more of the following:

GPU Infrastructure:

  • DCGM
  • BMC / Redfish
  • Firmware Lifecycle Management
  • Driver Lifecycle Management

Networking

  • InfiniBand
  • High-Speed Ethernet
  • NCCL
  • UFM

High-Performance Storage

  • Lustre
  • IBM Storage Scale (GPFS)
  • WEKA
  • VAST Data
  • Similar Enterprise Storage Platforms

Platform Experience

Hands-on experience with:

  • Kubernetes
  • Slurm
  • GPU Scheduling
  • Multi-Tenancy Architecture

Observability

  • Prometheus
  • Grafana
  • OpenTelemetry

Automation & Infrastructure as Code

  • Terraform
  • Ansible
  • Argo CD
  • Similar Automation Frameworks

Operating Systems & Scripting

  • Linux Administration
  • Python
  • Bash
  • Comparable scripting languages

Soft Skills

  • Strong root-cause analysis and troubleshooting skills.
  • Ability to communicate complex technical findings clearly.
  • Experience leading technical initiatives without direct authority.
  • Strong customer-facing communication skills.
  • Ability to manage multiple partner engagements simultaneously.

Preferred Qualifications

Candidates will stand out if they have:

  • Experience operating GPU cloud environments under production workloads.
  • Experience managing large-scale AI platforms or HPC environments.
  • Built or improved 24x7 operational support functions.
  • Experience with:

  • Observability platforms
  • Incident management
  • Problem management
  • On-call systems

Highly Desired

Hands-on experience with:

  • GB200 NVL72
  • GB300 NVL72
  • Spectrum-X
  • UFM
  • Base Command Manager
  • Mission Control
  • GPU Operators
  • Network Operators

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:15 min

Transitioning to hybrid clouds amid GPU scarcity

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

Videos

See all

Related articles

See all