> Markdown version of [/jobs/ext/1973640-principal-software-quality-engineer-gpu-machine-learning](https://www.wearedevelopers.com/jobs/ext/1973640-principal-software-quality-engineer-gpu-machine-learning). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Software Quality Engineer - GPU & Machine Learning - **Company:** Advanced Micro Devices, Inc. - **Location:** San Jose, CA, United States - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Automation of Tests, Intelligent Platform Management Interface, C++ (Programming Language), Computer Clusters, Code Coverage, Profiling, Software Quality, Code Review, Software Debugging, Distributed Systems, Firmware, Github, Python (Programming Language), Linux Kernel, Machine Learning, Open Source Technology, PCI Express, Software Architecture, Tensorflow, Verification and Validation (Software), Graphics Processing Unit (GPU), Performance Testing, Pytorch, Large Language Models, System-level Testing, Data Lakes, Production Code, Server Operating Systems & Platforms - **Published:** August 7, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/principal-software-quality-engineer-gpu-and-machine-learning-san-jose-ca-usa-58832226 ## About the Role _ with reproducible methodology and baselines Tasks * Software validation/QA experience; ability to lead complex systems validation * Expert Python for test automation; strong C++ for debugging/production code extension * Deep validation expertise in GPU software stacks, AI/ML frameworks, HPC runtimes, Linux kernel or GPU drivers, or distributed systems * Experience validating multi-GPU, multi-node server platforms with stress, soak, fault injection, and RAS testing * Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers * Contributions to validation/CI/test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar OSS * Experience leading adoption of agentic AI workflows (automated testing, AI-driven debugging, MCP, RAG-based engineering) * Experience operating large-scale GPU clusters (256+ GPUs) including fabric bring-up, health monitoring, diagnostics * Familiarity with AI training/inference and HPC benchmark methodologies S _ Experience with performance validation and profiling tools (rocprof, Omniperf, Nsight) * Familiarity with hardware lab automation (BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, topology-aware scheduling) * Experience with pre-silicon, emulation, and first-silicon accelerator bring-up Key requirements * ## Description Experteer Overview In this senior, IC-led role, you define and drive ROCm validation for compute workloads and server-class systems, shaping verification strategy and test infrastructure. You ensure ROCm readiness across multi-GPU and multi-node platforms, influencing releases and quality for hyperscalers, OEMs, and the open-source community. You tackle the hardest debugging and qualification challenges and mentor a team of validation engineers. This work directly supports scalable, production-ready ROCm deployments. Compensation / Benefits * Own end-to-end ROCm validation architecture across unit, integration, workload, performance, and system-level tests for multiple GPUs and server platforms * Define release-qualification gates and exit criteria for ROCm releases and drive the org to meet them * Architect and evolve test infrastructure including distributed test runners, CI fleets, hardware lab orchestration, data lakes, flaky-test detection, and pre-submit pipelines * Promote modern quality engineering practices (shift-left testing, test pyramids, contract testing, hermetic environments, deterministic reproducers, trunk validation) * Establish GitHub-based quality workflows (PR gating, required checks, code coverage, triage cadences) and manage cross-repo quality * Lead complex escalation debugging with cross-functional teams to root-cause failures and extend test coverage * Influence the roadmap with product, silicon, and software architecture to ensure validation readiness before tape-in milestones * Mentor and elevate senior staff and SDETs; contribute through design/code reviews and guidance * Represent ROCm validation externally in customer engagements, OEM qualification programs, and open-source initiatives * Lead system-level server node testing (multi-GPU topologies, PCIe/Infinity Fabric, BMC/IPMI, thermal/power, firmware interactions, multi-node fabric) * Drive compute workload validation and characterization (LLM training/inference, PyTorch, Triton, vLLM, JAX) with reproducible methodology and baselines Tasks * Software validation/QA experience; ability to lead complex systems validation * Expert Python for test automation; strong C++ for debugging/production code extension * Deep validation expertise in GPU software stacks, AI/ML frameworks, HPC runtimes, Linux kernel or GPU drivers, or distributed systems * Experience validating multi-GPU, multi-node server platforms with stress, soak, fault injection, and RAS testing * Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers * Contributions to validation/CI/test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar OSS * Experience leading adoption of agentic AI workflows (automated testing, AI-driven debugging, MCP, RAG-based engineering) * Experience operating large-scale GPU clusters (256+ GPUs) including fabric bring-up, health monitoring, diagnostics * Familiarity with AI training/inference and HPC benchmark methodologies * Experience with performance validation and profiling tools (rocprof, Omniperf, Nsight) * Familiarity with hardware lab automation (BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, topology-aware scheduling) * Experience with pre-silicon, emulation, and first-silicon accelerator bring-up Key requirements * ## Related Videos - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [Playing Pong on a shoulder press machine](https://www.wearedevelopers.com/videos/100140-playing-pong-on-a-shoulder-press-machine) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Agent Smith Gets Hardware: Autonomous IoT Hacking From Debug Port to Cloud API](https://www.wearedevelopers.com/videos/100258-agent-smith-gets-hardware-autonomous-iot-hacking-from-debug-port-to-cloud-api) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)