> Markdown version of [/jobs/ext/1994690-principal-software-quality-engineer-gpu-machine-learning](https://www.wearedevelopers.com/jobs/ext/1994690-principal-software-quality-engineer-gpu-machine-learning). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Software Quality Engineer - GPU & Machine Learning - **Company:** Advanced Micro Devices, Inc. - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $240,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Automation of Tests, Intelligent Platform Management Interface, C++ (Programming Language), Computer Clusters, Code Coverage, Profiling, Software Quality, Code Review, Nvidia CUDA, Computer Engineering, Software Debugging, Distributed Systems, Ethernet, Firmware, Github, InfiniBand, Python (Programming Language), Linux Kernel, Machine Learning, Regression Analysis, Node.Js, Open Source Technology, PCI Express, Recommender Systems, Tensorflow, Software Engineering, Graphics Processing Unit (GPU), Performance Testing, Pytorch, Large Language Models, System-level Testing, Information Technology, Production Code, Server Operating Systems & Platforms, SDET, Jenkins - **Published:** August 8, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/17864064?backUrl=%2Fcareer%2F17864064%2FPrincipal-Software-Quality-Engineer-Gpu-Machine-Learning-California-San-Jose ## About the Role * Software engineering experience in validation, SDET, or quality engineering, including experience leading complex systems validation. * Expert Python for test automation and infrastructure; strong C++ for debugging and extending production code. * Deep validation expertise in two or more of the following: + GPU software stacks (ROCm, CUDA, oneAPI, SYCL) + AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton, vLLM) + HPC runtimes and communication libraries (MPI, RCCL/NCCL, UCX, Libfabric) + Linux kernel, GPU drivers, or accelerator firmware + Distributed systems and large-scale cluster software * Experience validating multi-GPU, multi-node server platforms, including stress, soak, fault injection, and RAS testing. * Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers. * Contributions to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects. * Experience leading adoption of agentic AI workflows, including automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions. * Experience validating or operating large-scale GPU clusters (256+ GPUs), including fabric bring-up, health monitoring, and diagnostics. * Familiarity with AI training, inference, and HPC benchmark methodologies. * Experience with performance validation, profiling tools (rocprof, Omniperf, Nsight), and regression analysis. * Familiarity with hardware lab automation, including BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, and topology-aware scheduling. * Experience supporting validation for pre-silicon, emulation, and first-silicon accelerator bring-up., * BS/MS/PhDin Computer Science, Computer Engineering, orrelated discipline (or equivalent demonstrated experience). ## Description Weare seeking aPrincipal Software Engineertoserve as the senior technical leader forROCm software validationacrosscompute workloads and server-class systems. Inthis individual-contributor leadership role, you will definehowAMD provesROCm is ready to ship- from unit andcomponenttesting, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct GPU platforms. Youwill set the technical direction for validation strategy, build and evolve the test infrastructure thatgates everyROCm release, and personally drive the hardestdebugging, characterization, and qualification problems. Your work directly determines thequality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community runningROCm inproduction., Youwill set the technical direction for validation strategy, build and evolve the test infrastructure thatgates everyROCm release, and personally drive the hardestdebugging, characterization, and qualification problems. Your work directly determines thequality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community runningROCm inproduction., * Ownthe end-to-end validation architecture for ROCm - unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers - across multiple GPU generations and server platforms. * Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them. * Architect the test infrastructure - distributed test runners, GitHub Actions/ Jenkins / internal CI fleets, hardware lab orchestration, resultdatalakes, flaky-test detection, bisectionautomation, and self-servicedeveloper pre-submit pipelines. * Champion modern, agile quality engineering - shift-left testing, test pyramids, contract testing betweenlayers, hermetic test environments, deterministic reproducers, and continuous validation intrunk. * Set the bar for GitHub-based quality workflows - PR gatingpolicy, requiredchecks, code-coverage standards, bug-bashandtriage cadences, and disciplined issue management acrossROCm/*repositories and partner upstream projects. * Lead complex escalation debug - partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage. * Influence the roadmap- work with product management, silicon, platform, and softwarearchitecture to ensure validation readiness fornext-generation Instinct GPUs and serverplatformsbefore tape-inmilestones and silicon arrival. * Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through designreview, code review, and written guidance. * Represent ROCm validation externally - strategic customer engagements, OEM qualification programs, and open-source community quality initiatives. * Lead system-level testing for server nodes- multi-GPU topologies, PCIe/InfinityFabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric(Ethernet/InfiniBand/UALink) bring-up andvalidation.Drive compute workload validation and characterization- LLM training andinference(PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks- establishing reproducible methodology, baselines, and regression tracking., AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here. ## Related Videos - [Coffee with Developers - Stephen Jones - NVIDIA](https://www.wearedevelopers.com/videos/1303-coffee-with-developers-stephen-jones-nvidia) - [The Road to MLOps: How Verivox Transitioned to AWS](https://www.wearedevelopers.com/videos/1050-the-road-to-mlops-how-verivox-transitioned-to-aws) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Accelerating Python on GPUs](https://www.wearedevelopers.com/videos/859-accelerating-python-on-gpus) - [Our GitOps approach for deploying an Identity Provider and an API Gateway in a SaaS company](https://www.wearedevelopers.com/videos/776-our-gitops-approach-for-deploying-an-identity-provider-and-an-api-gateway-in-a-saas-company) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence)