Principal Software Quality Engineer - GPU & Machine Learning

Advanced Micro Devices, Inc.
San Jose, CA, United States
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence Automation of Tests Intelligent Platform Management Interface C++ (Programming Language) Computer Clusters Code Coverage Profiling Software Quality Code Review Software Debugging Distributed Systems Firmware
+17 more
Github Python (Programming Language) Linux Kernel Machine Learning Open Source Technology PCI Express Software Architecture Tensorflow Verification and Validation (Software) Graphics Processing Unit (GPU) Performance Testing Pytorch Large Language Models System-level Testing Data Lakes Production Code Server Operating Systems & Platforms

Job description

Experteer Overview In this senior, IC-led role, you define and drive ROCm validation for compute workloads and server-class systems, shaping verification strategy and test infrastructure. You ensure ROCm readiness across multi-GPU and multi-node platforms, influencing releases and quality for hyperscalers, OEMs, and the open-source community. You tackle the hardest debugging and qualification challenges and mentor a team of validation engineers. This work directly supports scalable, production-ready ROCm deployments. Compensation / Benefits * Own end-to-end ROCm validation architecture across unit, integration, workload, performance, and system-level tests for multiple GPUs and server platforms * Define release-qualification gates and exit criteria for ROCm releases and drive the org to meet them * Architect and evolve test infrastructure including distributed test runners, CI fleets, hardware lab orchestration, data lakes, flaky-test detection, and pre-submit pipelines * Promote modern quality engineering practices (shift-left testing, test pyramids, contract testing, hermetic environments, deterministic reproducers, trunk validation) * Establish GitHub-based quality workflows (PR gating, required checks, code coverage, triage cadences) and manage cross-repo quality * Lead complex escalation debugging with cross-functional teams to root-cause failures and extend test coverage * Influence the roadmap with product, silicon, and software architecture to ensure validation readiness before tape-in milestones * Mentor and elevate senior staff and SDETs; contribute through design/code reviews and guidance * Represent ROCm validation externally in customer engagements, OEM qualification programs, and open-source initiatives * Lead system-level server node testing (multi-GPU topologies, PCIe/Infinity Fabric, BMC/IPMI, thermal/power, firmware interactions, multi-node fabric) * Drive compute workload validation and characterization (LLM training/inference, PyTorch, Triton, vLLM, JAX) with reproducible methodology and baselines Tasks * Software validation/QA experience; ability to lead complex systems validation * Expert Python for test automation; strong C++ for debugging/production code extension * Deep validation expertise in GPU software stacks, AI/ML frameworks, HPC runtimes, Linux kernel or GPU drivers, or distributed systems * Experience validating multi-GPU, multi-node server platforms with stress, soak, fault injection, and RAS testing * Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers * Contributions to validation/CI/test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar OSS * Experience leading adoption of agentic AI workflows (automated testing, AI-driven debugging, MCP, RAG-based engineering) * Experience operating large-scale GPU clusters (256+ GPUs) including fabric bring-up, health monitoring, diagnostics * Familiarity with AI training/inference and HPC benchmark methodologies * Experience with performance validation and profiling tools (rocprof, Omniperf, Nsight) * Familiarity with hardware lab automation (BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, topology-aware scheduling) * Experience with pre-silicon, emulation, and first-silicon accelerator bring-up Key requirements *

Requirements

_ with reproducible methodology and baselines Tasks * Software validation/QA experience; ability to lead complex systems validation * Expert Python for test automation; strong C++ for debugging/production code extension * Deep validation expertise in GPU software stacks, AI/ML frameworks, HPC runtimes, Linux kernel or GPU drivers, or distributed systems * Experience validating multi-GPU, multi-node server platforms with stress, soak, fault injection, and RAS testing * Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers * Contributions to validation/CI/test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar OSS * Experience leading adoption of agentic AI workflows (automated testing, AI-driven debugging, MCP, RAG-based engineering) * Experience operating large-scale GPU clusters (256+ GPUs) including fabric bring-up, health monitoring, diagnostics * Familiarity with AI training/inference and HPC benchmark methodologies S _ Experience with performance validation and profiling tools (rocprof, Omniperf, Nsight) * Familiarity with hardware lab automation (BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, topology-aware scheduling) * Experience with pre-silicon, emulation, and first-silicon accelerator bring-up Key requirements *

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com

Good distractions

Talks and stories from around this role β€” technically off-topic, practically not.

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters Β· WWC 2023

2:35 min

Preventing remote code execution in PyTorch models

BalΓ‘zs Kiss Β· WWC 2023

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac Β· WWC 2021

1:34 min

Profiling and debugging GPU code with Nsight developer tools

Paul Graham Paul Graham Β· WWC 2025

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle Β· Coffee With Developers

3:20 min

Certifying safety-critical automotive software and deploying progressive testing strategies

David Romić · WWC 2023

Videos

See all

Related articles

See all