Skip to content

Session

Trust, But Verify: Continuous GPU Validation at Scale

with Kyle Bell

About This Session

AI infrastructure has long been dominated by a single ecosystem, but the model is changing. In this session, Kyle Bell from TensorWave explores how to build & operate large-scale AI clusters on AMD Instinct GPUs using Kubernetes as the orchestration backbone. Attendees will learn about: AMD-specific technologies like ROCm, GPU Operator, & RVS (ROCm Validation Suite). Kubernetes integration patterns for AI scheduling, node triage, & GPU telemetry. Slurm-on-Kubernetes for HPC-style orchestration, & how it impacts resource management and observability. Custom convergence testing & GPU validation pipelines that ensure reliability at scale. This talk provides an open-source roadmap for operators, ML engineers, and platform teams who want to move beyond vendor lock-in while maintaining reliability, performance, & observability at scale. Viewers will walk away with practical guidance on designing Kubernetes clusters purpose-built for AMD Instinct GPUs, integrating ROCm into existing cloud-native toolchains, applying reliability engineering patterns for AI workloads, & building open sustainable infrastructure that contributes to a more diverse AI hardware ecosystem.

Topics

  • Multi-Cloud