Senior Director, System Software Engineering - DGX...

NVIDIA Ltd.
Santa Clara, CA, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Computing Platforms Cloud Computing Cloud Engineering Software Quality Computer Engineering DevOps Distributed Systems Open Source Technology Performance Tuning Software Engineering Software Systems
+4 more
System Software AI Infrastructure HybridCloud Information Technology

Job description

NVIDIA is seeking a Senior Director, System Software Engineering, to lead strategy and execution for capacity management in DGX Cloud, building the capacity foundation for NVIDIA’s internal AI research clusters. This leader will shape the roadmap for scalable system software that automates GPU management at scale, drive execution across teams and functions, and partner closely with architecture, security, product, and developer platform leaders to deliver reliable, high-performance software that powers the next generation of accelerated computing. The ideal candidate combines deep systems expertise with strong organizational leadership, technical judgment, and builds teams that deliver sophisticated platform software at scale.

What you’ll be doing:

  • Define and drive the system software strategy for capacity management and automation for DGX Cloud’s GPU cloud platforms, aligning long-range technical direction with business and product priorities.

  • Lead engineering leaders responsible for core platform capabilities such as runtime software, host and cluster management, provisioning, observability, reliability, security, and performance optimization.

  • Build a strong execution model across planning, architecture reviews, release readiness, quality, and operational excellence for software delivered across on-prem and cloud environments.

  • Partner closely with security, DevOps, research, and product organizations to translate platform requirements into scalable software roadmaps and high-quality releases.

  • Establish measurable goals for engineering efficiency, service reliability, software quality, and customer impact, using data to continuously improve delivery and operations.

Requirements

  • BS, MS, or PhD in Computer Science, Computer Engineering, or a related technical field, or equivalent experience.

  • 16+ overall years of relevant management experience in system software, platform software, or distributed systems engineering, 7+ years of significant leadership experience leading engineering organizations.

  • Deep technical expertise in operating systems, distributed systems, platform architecture, cloud infrastructure, or large-scale systems software.

  • Demonstrated experience leading delivery of complex software platforms spanning reliability, performance, scalability, security, and observability.

  • Strong record of leadership and influence across engineering, product, program management, and executives.

  • Demonstrated success building and leading high-performing teams, developing leaders, and scaling organizations through growth and change.

  • Excellent technical communication and decision-making, with the ability to connect architecture choices to business outcomes.

  • Demonstrated experience with industry-leading AI tools that help engineers and engineering leaders work more efficiently.

Ways to stand out from the crowd:

  • Experience with AI infrastructure, accelerated computing, GPU-optimized software stacks, or large-scale training and inference environments.

  • Experience leading platform software for cloud-native or hybrid-cloud deployments.

  • Track record of driving architectural simplification and operational excellence across large, complex engineering portfolios.

  • Experience partnering with open-source communities and ecosystem partners on platform adoption and enablement.

Benefits & conditions

Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 384,000 USD - 575,000 USD.

You will also be eligible for equity and benefits (https://www.nvidia.com/en-us/benefits/) .

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:27 min

Introduction to WebAssembly in a cloud computing context

Edo Edo · WWC 2024

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · WWC 2021

2:51 min

Alibaba Cloud developer resources and cloud computing training

Cheng Zhang · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all