GPU Solutions Architect - Cloud
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+15 more
Requirements
- BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Mathematics, Physics, or related discipline (or equivalent experience).
-
8+ years of experience in:
- Production Infrastructure
- Cloud Engineering
- Solutions Architecture
- Site Reliability Engineering (SRE)
- HPC Environments
- Similar technical disciplines OR
5+ years of exceptional specialist-level experience supporting large-scale GPU or AI infrastructure.
Technical Expertise: Experience building, operating, and optimizing distributed infrastructure in production environments. Deep hands-on expertise in one or more of the following:
GPU Infrastructure:
- DCGM
- BMC / Redfish
- Firmware Lifecycle Management
- Driver Lifecycle Management
Networking
- InfiniBand
- High-Speed Ethernet
- NCCL
- UFM
High-Performance Storage
- Lustre
- IBM Storage Scale (GPFS)
- WEKA
- VAST Data
- Similar Enterprise Storage Platforms
Platform Experience
Hands-on experience with:
- Kubernetes
- Slurm
- GPU Scheduling
- Multi-Tenancy Architecture
Observability
- Prometheus
- Grafana
- OpenTelemetry
Automation & Infrastructure as Code
- Terraform
- Ansible
- Argo CD
- Similar Automation Frameworks
Operating Systems & Scripting
- Linux Administration
- Python
- Bash
- Comparable scripting languages
Soft Skills
- Strong root-cause analysis and troubleshooting skills.
- Ability to communicate complex technical findings clearly.
- Experience leading technical initiatives without direct authority.
- Strong customer-facing communication skills.
- Ability to manage multiple partner engagements simultaneously.
Preferred Qualifications
Candidates will stand out if they have:
- Experience operating GPU cloud environments under production workloads.
- Experience managing large-scale AI platforms or HPC environments.
- Built or improved 24x7 operational support functions.
-
Experience with:
- Observability platforms
- Incident management
- Problem management
- On-call systems
Highly Desired
Hands-on experience with:
- GB200 NVL72
- GB300 NVL72
- Spectrum-X
- UFM
- Base Command Manager
- Mission Control
- GPU Operators
- Network Operators
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
Highest Paying Tech Companies for Developers
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again