Senior Software Engineer, At Scale Compute Analysis

NVIDIA Corporation
Santa Clara, CA, United States
7 days ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$152,000.0 - $241,500.0
Working hours
Regular working hours

Tech stack

JavaScript (Programming Language) Artificial Intelligence Systems Engineering Big Data Computer Graphics Software Debugging Linux Elasticsearch Python (Programming Language) Machine Learning Tensorflow Pytorch
+3 more
Grafana Information Technology Splunk

Job description

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology-and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

Join a team that analyzes large-scale datacenter workloads on GPU-accelerated clusters. You will turn telemetry and workload data into clear findings and visuals. You will partner with OS, container, GPU, and systems engineers. When useful, you will apply machine learning and deep learning techniques for categorization and forecasting. These will be coordinated into tools the team actually uses.

What you’ll be doing:

  • Analyze large-scale workloads and infrastructure signals to find application and platform improvement opportunities.
  • Work with high-dimensional data: spot trends, tie changes to known events, summarize conclusions, and communicate results to engineers and leadership.
  • Partner with the team to clarify questions, scope analyses, and document methods so others can extend your work.
  • Build and maintain practical visualizations and lightweight implementations (e.g. ML/DL for classification/prediction) inside existing software workflows.

Requirements

  • 5+ years analyzing complex datasets, debugging data issues, and communicating trends clearly.
  • BS or MS in Engineering, Mathematics, Physics, Computer Science, or equivalent experience.
  • Strong Python and JavaScript;
  • Comfortable being responsible for an analysis end-to-end.
  • Hands-on use of telemetry / observability stacks (e.g. Grafana, Elasticsearch, Splunk).
  • Shown grasp of core ML concepts; quick learner; strong analytical and problem-solving skills.
  • Collaboration and communication.

Ways to stand out from the crowd:

  • TensorFlow or PyTorch
  • Linux and HPC / large-scale or performance-sensitive environments
  • Experience visualizing high-dimensional problems
  • Diligent, action-biased analysis style

Benefits & conditions

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you’re creative and autonomous, we want to hear from you!

NVIDIA also offers a comprehensive benefits package. We provide health care coverage, dental and vision, 401(K), including company matching and after tax contributions, Employee Stock Purchase Program (ESPP), Employee Assistance Program (EAP), company paid holidays, paid sick leave, vacation leave, professional time off, life and disability protection.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD.

You will also be eligible for equity and benefits.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:08 min

History and scale of NVIDIA GPU computing

Paul Graham Paul Graham · LIVE

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all