Senior Software QA Test Development Engineer - Diagnostics
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+25 more
Job description
Experteer Overview As a Platform SWQA Engineer at NVIDIA, you will design and execute test plans for HGX/DGX/MGX server platforms across OS, firmware, and CUDA software. You’ll install, test, and automate validation, drive root-cause analysis for reliability issues, and build end-to-end automation for server and OS testing. You’ll review partner results and push for targeted reliability tests, while collaborating in an agile, quality-driven environment. This role combines hardware expertise with software QA to ensure scalable, high-quality AI hardware platforms. Compensation / Benefits * Develop and execute platform test plans for HGX/DGX/MGX servers across OS, firmware, and CUDA software * Install and validate systems OS, server firmware, and software stack * Lead root-cause analysis of reliability/validation failures and implement mitigations * Build and maintain automation front-end and back-end tests for server and OS level * Review partner/supplier test results and prescribe additional reliability testing * Work in an agile software development team with high production quality standards * Manage bug lifecycle and collaborate with cross-group teams to drive solutions Tasks * Bachelor’s degree in STEM or equivalent experience * 5+ years of experience or master’s degree * Proven OS and server automation, CI/CD, and DevOps experience using Python, Shell, Ansible, Jenkins, C/C++, Java, JavaScript * Strong Linux troubleshooting in bare-metal and virtualization environments (KVM/VMware/Hyper-V) * Experience with AI tools/frameworks (TensorFlow, PyTorch) and NLP/LLM benchmarking * Experience creating test plans, test cases, and automation for AI development tools * Knowledge of FW, BMC/OpenBMC, network protocols, enterprise storage, PCIe, IO devices, CPU/memory, ACPI, UEFI, Redfish; virtualization and hardware interfaces is a plus * Experience with GitHub/GitLab/Gerrit, PXE, SLURM, Stack/Kubernetes/Docker is a plus Key requirements * equity * benefits
Requirements
across reliability testing * Work in an agile software development team with high production quality standards * Manage bug lifecycle and collaborate with cross-group teams to drive solutions Tasks * Bachelor’s degree in STEM or equivalent experience * 5+ years of experience or master’s degree * Proven OS and server automation, CI/CD, and DevOps experience using Python, Shell, Ansible, Jenkins, C/C++, Java, JavaScript * Strong Linux troubleshooting in bare-metal and virtualization environments (KVM/VMware/Hyper-V) * Experience with AI tools/frameworks (TensorFlow, PyTorch) and NLP/LLM benchmarking * Experience creating test plans, test cases, and automation for AI development tools * Knowledge of FW, BMC/OpenBMC, network protocols, enterprise storage, PCIe, IO devices, CPU/memory, ACPI, UEFI, Redfish; virtualization and hardware interfaces is a plus * Experience with GitHub/GitLab/Gerrit, PXE, SLURM, Stack/Kubernetes/Docker is a plus Key requirements * equity * benefits
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
The 8 Best Code Testing Tools
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
Dev Digest 120 - Apple and peers