Senior System Software Engineer, Enterprise MODS

NVIDIA Ltd.
Santa Clara, CA, United States
16 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$224,000.0
Working hours
Regular working hours
Job source

Tech stack

C (Programming Language) Microsoft Windows Artificial Intelligence ARM Architecture BIOS C++ (Programming Language) Linux Ethernet Firmware InfiniBand Python (Programming Language) PCI Express
+7 more
Systems Development Life Cycle System Software Extensible Firmware Interface Cloud Platform System High Performance Computing Information Technology Programming Languages

Job description

NVIDIA is at the forefront of AI, HPC, and visualization. Our diagnostics are the nervous system of our platforms-ensuring reliability, performance, and innovation at scale. If you’re a creative, driven architect ready to shape the future of diagnostics, we want to hear from you.

Requirements

  • Proven experience architecting diagnostics for complex server systems, especially at the SW/HW interface.

  • Deep systems knowledge: x86/ARM architectures, Linux/Windows OS internals, firmware (UEFI/BIOS), BMC, and platform security.

  • Ability to weigh tradeoffs in system development and drive the most optimum solutions with customers and multi-disciplinary teams

  • Expertise in programming languages like C, C++, and Python for tool development and automation.

  • Familiarity with high-speed interconnects such as PCIe, Infiniband, NVLink, and Ethernet.

  • Strong communication skills to engage with technical and executive team.

  • BS/MS or equivalent experience in Computer Science, Electrical Engineering, or related field.

  • 12+ years of engineering experience in diagnostics, embedded systems, or cloud platforms.

Ways to stand out from the crowd:

  • Experience driving diagnostics across rack-level or cluster-level deployments.

  • Background in cloud-scale infrastructure and partner engagement.

  • Demonstrated success in influencing product direction and vendor roadmaps.

  • Passion for mentoring and building high-performing teams.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 5, and 272,000 USD - 431,250 USD for Level 6.

About the company

NVIDIA (Santa Clara, CA)

At NVIDIA, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work.

The data center platforms like GB200 NVL72 by NVIDIA are redefining AI, HPC, and cloud computing. To accommodate leading workloads globally, our diagnostic systems need to evolve across diverse hardware technologies. We’re in search of a visionary technical leader to engineer and propel innovation in diagnostics for NVIDIA’s partner ecosystem. This role is essential in crafting how we validate, debug, and optimize complex server platforms across ODM factories, Cloud Service Provider (CSP) deployments, and field operations.

What You’ll Be Doing:

  • Develop diagnostic systems for NVIDIA data center platforms, which involve hardware and software tools to develop the worst case stress workloads for CPUs, GPUs, memory, storage, and interconnects.

  • Lead platform bring-up and integration, ensuring diagnostics are embedded early and effectively across the server lifecycle.

  • Drive hardware validation strategy in collaboration with architecture and hardware teams, crafting robust validation plans for new server generations.

  • Analyze root causes of complex failures, acting as a Level 2 engineering contact for critical issues and offering scalable solutions across the stack.

  • Develop diagnostics software to ensure quality and performance at scale across ODM and partner production lines.

  • Mentor and grow engineering teams, providing technical leadership and encouraging a culture of innovation and excellence.

  • Influence the long-term strategy by developing diagnostic architecture and roadmaps for the upcoming products of NVIDIA and its partners.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

41 sec

Massive client data loss and bio-digital storage

Chris Heilmann +1 · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · WWC 2025

4:52 min

Connecting namespaces with local virtual ethernet pairs

Oliver Seitz Oliver Seitz · WWC 2025

1:12 min

Addressing the competitive landscape of specialized hardware demands

Hazal Mestci +1 · Coffee With Developers

1:19 min

Advancing autonomous driving capabilities with specialized software talent

Katrin Lehmann Katrin Lehmann +1 · Coffee With Developers

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

Videos

See all

Related articles

See all