Senior AI GPU Deployment Engineer
Role details
Job location
Tech stack
Job description
We are seeking an experienced Senior AI GPU Deployment Engineer to plan, deploy, and operationalize large-scale GPU AI infrastructure environments. This role delivers production-grade GPU clusters supporting AI training, inference, and high-performance computing workloads in our hyperscale data centers., We're looking for a Senior AI GPU Deployment Engineer to join our Cloud Ops team in the United States. This person will share our company values and play an important role in supporting our continued growth.
What You Will Do
- Deploy integrate and validate multi-rack GPU-based compute platform deployments
- Deploy fabric configuration engines (Subnet Manager), observability platforms (UFM) and validate interconnect and fabric performance (nccl)
- Collaborate with network engineering team on topology implementation and optimization and storage engineering team on deployment and integration of high-performance storage environments supporting AI workloads (e.g. VAST Data)
- Configure settings and manage firmware updates for GPUs, NICs, BMC, BIOS and other components across large-scale clusters
- Contribute to infrastructure-as-code automation development for cluster provisioning and lifecycle management
- Contribute to improving and documenting repeatable deployment methodologies and scalable operational standards
- Query and analyze deployment outcomes using SQL for diagnostics and operational reporting
Requirements
The ideal candidate brings deep technical expertise in GPU infrastructure, network fabrics, storage, automation, and Linux systems administration, combined with strong execution and troubleshooting skills., * Bachelor's degree in Computer Science, Engineering, IT, or related field (or equivalent experience)
- 5+ years of infrastructure engineering or datacenter deployment experience
- 3+ years deploying large-scale AI, HPC, or GPU infrastructure
- Hands-on experience deploying and operating large GPU clusters in enterprise or hyperscale environments
- Strong expertise with:
- GPU architectures
- InfiniBand (NDR/XDR) and Ethernet GPU fabrics (Spectrum-X)
- NVLink, NVSwitch, and GPU-direct technologies
- Canonical MaaS and automated provisioning systems
- VAST Data or similar high-performance storage platforms
- Linux systems administration for HPC/AI workloads
- Infrastructure-as-Code and configuration management (Ansible)
- Python, Shell, and SQL for infrastructure automation and diagnostics
- Strong understanding of:
- RDMA, RoCE, and lossless Ethernet fabrics
- Cluster automation, observability, and lifecycle management
Benefits & conditions
Pulled from the full job description
- 401(k)
- Health insurance
- Vision insurance
- Dental insurance
- Pension plan