Senior Software Development Engineer in Test

NVIDIA Corporation
Santa Clara, CA, United States
1 day ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Compensation
$168,000.0 - $270,250.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Automation of Tests Microsoft Azure Cloud Computing Cloud Engineering Code Coverage Continuous Integration Data Centers Dynamic Host Configuration Protocol Software Debugging Linux
+27 more
Distributed Systems Domain Name System (DNS) Monitoring of Systems Hypertext Transfer Protocols (HTTP) Software Reliability Testing Cloud Services Ansible Prometheus Selenium Software Engineering Web Applications Google Cloud Enterprise Software Applications Cloud Platform System High Performance Computing Grafana Gitlab Kubernetes Information Technology Playwright Slurm Cloudwatch Terraform Oracle Cloud Infrastructure Docker SDET Jenkins

Job description

  • Work with development teams on test plans for all layers of SW stack for cloud infrastructure, execution, reviews, failure analysis and assessing overall quality and risk. Work with customer PMs on software issues including technical feedback from OEMs and CSPs. Develop key benchmarks to track execution and deploy process improvements to improve efficiency
  • Leverage AI skills to expedite the test scope, test plan, execution and automation workflows.
  • Lead NVIDIA Cloud and Data Center bring up activities which will involve validation, reporting, working with engineering to debug issues, providing design input at times, adding coverage in different areas.
  • Design, develop and maintain CI/CD pipelines for continuous testing in cloud environments when needed.
  • Perform performance, scalability, and reliability testing of cloud services.
  • Implement and maintain test environments in cloud platforms such as AWS, Azure, or Google Cloud.
  • Supervise the infrastructure to alert on significant events, ensuring the highest level of system performance and reliability.
  • Work with various different partner teams to ensure availability of clusters to test on and take the lead in resolve all issues.
  • Working with teams to ensure quality of the cloud products getting delivered focusing on critical areas like security, storage, workloads, performance on latest SW and FW components.

Requirements

We are seeking a highly skilled and hard-working Senior Test Developer / test engineer to join our multifaceted Enterprise Software QA team. This role offers an outstanding opportunity to leave your mark on the design, construction, optimization and testing of large-scale infrastructure for various foundational NVIDIA unified cloud services and data center offerings. If you are a dedicated engineer with strong expertise in cloud infrastructure and distributed systems and want to apply your skills with AI tools, this role could fit you perfectly. You will thrive in an exciting, innovative environment., * A Master’s or Ph.D. in Computer Science or a related field, or equivalent experience.

  • Experience with AI development tools used in creating test cases, automating test cases, code coverage, triaging.
  • 8+ years of hands-on experience in cluster management and related tools, including Docker Containers, Slurm, Kubernetes, and Ansible.
  • 2+ years strong experience with cloud infrastructure platforms like AWS, Azure, Google, OCI Cloud.
  • Hands-on experience with network, storage, security, cluster configuration and debugging, cloud infrastructure management tools like terraform, ansible.
  • Expertise in administering, operating, and configuring Kubernetes.
  • Experience in CI/CD tools such as Gitlab and Jenkins and the GitOps model.
  • Proficiency in various monitoring tools :Prometheus, Grafana, Cloudwatch, and Thanos.
  • Proficiency in debugging issues involving networks, DHCP, DNS, HTTP, Linux, and containers.

Ways to Stand Out from the Crowd:

  • Familiarity with “Base Command Manager” for managing and monitoring high performance computing.
  • Experience in writing automation for web application using tools like selenium, playwright.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 270,250 USD.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:06 min

Elevating the QA engineering role for complex challenges

Ondřej Gróf Ondřej Gróf · World Congress 2026 Europe

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all