Senior Machine Learning Engineer

NVIDIA Ltd.
Santa Clara, CA, United States
8 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Compensation
$184,000.0 - $287,500.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Cloud Computing Distributed Systems Python (Programming Language) Machine Learning Tensorflow Software Engineering Graphics Processing Unit (GPU) Pytorch Large Language Models Kubernetes
+2 more
Optimization Algorithms Operational Systems

Job description

As a Senior Machine Learning Engineer at NVIDIA, you will build the machine learning brain that keeps NVIDIA’s global DGX Cloud healthy, efficient and ready for the next waves of AI breakthroughs. DGX Cloud fuses NVIDIA GPUs, NVLink networking and the full AI software stack into elastic infrastructure powering large language models, drug discovery, autonomous driving and climate science. Your models will turn billions of telemetry signals into predictive insight. This frees customers to innovate while our platform runs smarter.

What you’ll be doing:

  • Ground breaking and developing innovative machine learning algorithms and models that propel our AI products.
  • Build production models for anomaly detection, predictive maintenance and usage optimization.
  • Develop tools surfacing real time telemetry, efficiency metrics and long term trends.
  • Develop forecasting and simulation models for global scale planning.
  • Analyzing complex datasets to determine the best approach for model training and optimization.
  • Translate findings into clear engineering actions with infrastructure, operations and product teams.
  • Participating in cross-functional projects to integrate machine learning capabilities into various NVIDIA products.

Requirements

  • Master’s degree or PhD in Mathematics, Statistics, Machine Learning or related quantitative field (or equivalent experience).
  • 8+ years experience applying Machine Learning to operational systems.
  • Proven track record of building and deploying Machine Learning models in production environments.
  • Experience with time series analysis and optimization algorithms.
  • Familiarity with distributed systems and cloud platforms such as AWS and Kubernetes.
  • Strong software engineering skills and proficiency in Python.
  • Effective verbal/written communication, and technical presentation skills.
  • Experience with machine learning frameworks such as TensorFlow, PyTorch, or similar.
  • A track record of delivering high-impact projects to compete in a fast-paced environment.

Ways to stand out from the crowd:

  • Experience solving capacity planning problems.
  • Deep understanding of GPU performance metrics.
  • Familiarity with prometheus and PromQL.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jofdav.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

1:39 min

Fundamentals of tensors and the TensorFlow library

Håkan Silfvernagel · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · WWC Europe 2026

2:15 min

Open-source community and machine learning frameworks

Gian Marco Iodice Gian Marco Iodice · WWC 2025

Videos

See all

Related articles

See all