Principal Software Quality Engineer - GPU & Machine Learning
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+17 more
Job description
Experteer Overview In this senior, IC-led role, you define and drive ROCm validation for compute workloads and server-class systems, shaping verification strategy and test infrastructure. You ensure ROCm readiness across multi-GPU and multi-node platforms, influencing releases and quality for hyperscalers, OEMs, and the open-source community. You tackle the hardest debugging and qualification challenges and mentor a team of validation engineers. This work directly supports scalable, production-ready ROCm deployments. Compensation / Benefits * Own end-to-end ROCm validation architecture across unit, integration, workload, performance, and system-level tests for multiple GPUs and server platforms * Define release-qualification gates and exit criteria for ROCm releases and drive the org to meet them * Architect and evolve test infrastructure including distributed test runners, CI fleets, hardware lab orchestration, data lakes, flaky-test detection, and pre-submit pipelines * Promote modern quality engineering practices (shift-left testing, test pyramids, contract testing, hermetic environments, deterministic reproducers, trunk validation) * Establish GitHub-based quality workflows (PR gating, required checks, code coverage, triage cadences) and manage cross-repo quality * Lead complex escalation debugging with cross-functional teams to root-cause failures and extend test coverage * Influence the roadmap with product, silicon, and software architecture to ensure validation readiness before tape-in milestones * Mentor and elevate senior staff and SDETs; contribute through design/code reviews and guidance * Represent ROCm validation externally in customer engagements, OEM qualification programs, and open-source initiatives * Lead system-level server node testing (multi-GPU topologies, PCIe/Infinity Fabric, BMC/IPMI, thermal/power, firmware interactions, multi-node fabric) * Drive compute workload validation and characterization (LLM training/inference, PyTorch, Triton, vLLM, JAX) with reproducible methodology and baselines Tasks * Software validation/QA experience; ability to lead complex systems validation * Expert Python for test automation; strong C++ for debugging/production code extension * Deep validation expertise in GPU software stacks, AI/ML frameworks, HPC runtimes, Linux kernel or GPU drivers, or distributed systems * Experience validating multi-GPU, multi-node server platforms with stress, soak, fault injection, and RAS testing * Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers * Contributions to validation/CI/test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar OSS * Experience leading adoption of agentic AI workflows (automated testing, AI-driven debugging, MCP, RAG-based engineering) * Experience operating large-scale GPU clusters (256+ GPUs) including fabric bring-up, health monitoring, diagnostics * Familiarity with AI training/inference and HPC benchmark methodologies * Experience with performance validation and profiling tools (rocprof, Omniperf, Nsight) * Familiarity with hardware lab automation (BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, topology-aware scheduling) * Experience with pre-silicon, emulation, and first-silicon accelerator bring-up Key requirements *
Requirements
_ with reproducible methodology and baselines Tasks * Software validation/QA experience; ability to lead complex systems validation * Expert Python for test automation; strong C++ for debugging/production code extension * Deep validation expertise in GPU software stacks, AI/ML frameworks, HPC runtimes, Linux kernel or GPU drivers, or distributed systems * Experience validating multi-GPU, multi-node server platforms with stress, soak, fault injection, and RAS testing * Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers * Contributions to validation/CI/test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar OSS * Experience leading adoption of agentic AI workflows (automated testing, AI-driven debugging, MCP, RAG-based engineering) * Experience operating large-scale GPU clusters (256+ GPUs) including fabric bring-up, health monitoring, diagnostics * Familiarity with AI training/inference and HPC benchmark methodologies S _ Experience with performance validation and profiling tools (rocprof, Omniperf, Nsight) * Familiarity with hardware lab automation (BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, topology-aware scheduling) * Experience with pre-silicon, emulation, and first-silicon accelerator bring-up Key requirements *
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on us.experteer.comGood distractions
Talks and stories from around this role β technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Dev Digest 120 - Apple and peers
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
Dev Digest 121 - AI goes offline
Highest Paying Tech Companies for Developers