Software Platform Support Engineer - GPU Cloud

NVIDIA Corporation
Santa Clara, CA, United States
26 days ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$108,000.0 - $172,500.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Computing Platforms Microsoft Azure Cloud Computing Databases Linux DevOps Distributed Computing Environment Reliability Engineering Software Systems Google Cloud Data Storage Technologies
+5 more
Kubernetes Information Technology Slurm Machine Learning Operations Oracle Cloud Infrastructure

Job description

The NVIDIA DGX Cloud organization is looking for passionate software support engineers to partner closely with our internal customers to support them on our internal platforms. This partnership requires you to gain a deep understanding of the customer needs, how their application(s) work, assist them in troubleshooting issues, and create documentation to make it easier for users to troubleshoot issues themselves in an ambiguous / fast-moving environment. The support you provide will help our users have a better experience and help shape our platform.

We expect you to have knowledge of supporting cloud-based deployments across compute, storage and networking environments.

What will you be doing:

  • Coordinate with multiple internal teams to provide Tier 1 support for complex cloud platforms
  • Define and improve operational workflows (runbooks, escalation paths, support processes)
  • Triage/investigate root cause of customer issues and escalate as needed
  • File bugs and report issues while working closely with the Site Reliability team
  • Build tooling to improve customer support process and visibility
  • Deeply understand user workloads and use cases
  • Partner with multiple internal teams to give feedback to engineering teams and develop solutions to aid in their success
  • Be part of an on call rotation to support production systems

Requirements

  • BS/MS degree in Computer science or related areas (or equivalent experience)
  • 5+ yrs of experience with supporting distributed software systems, supporting end-user software platforms, and experience with Linux
  • Experience with Kubernetes, AWS, Azure, OCI, and GCP
  • Background of Infrastructure, Networking, Storage, and DevOps scripting/tooling
  • Understanding of data storage technologies (databases, file, block, blob)
  • Customer Service/Support Experience
  • Willingness to work up and down the stack as well as across multiple teams
  • Strong skills in troubleshooting and Communication

Ways to stand out from the crowd:

  • Experience with MLOps workflows or ML infrastructure
  • Familiarity with GPU workloads or distributed training systems
  • SLURM or HPC previous experience
  • Strong drive to work with internal customers and make them successful
  • A drive to improve process with strong organizational skills

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 108,000 USD - 172,500 USD.

About the company

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:15 min

Transitioning to hybrid clouds amid GPU scarcity

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:22 min

Analyzing differences between mobile and traditional backend DevOps

Mete Baydar Mete Baydar · World Congress 2025

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:27 min

Defining DevOps through its historical origins and foundational texts

Sonal Patil · LIVE

1:51 min

Managing GPU quotas and multi-tenancy with Kueue

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

Videos

See all

Related articles

See all