Principal Systems Software Engineer, LPU

NVIDIA Corporation
Santa Clara, CA, United States
7 days ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$272,000.0
Working hours
Regular working hours

Tech stack

Board Bringup Abstraction Layers Application Programming Interfaces (APIs) Computer Programming Computer Engineering Continuous Integration Data Centers Extract Transform Load (ETL) Linux Distributed Systems Firmware Real-Time Operating Systems
+5 more
Reduced Instruction Set Computing Software Engineering System Software Kubernetes Api Management

Job description

We are now looking for a Principal Software Engineer for LPX System Software! NVIDIA’s LPX System Software team builds the foundational software that turns a novel deterministic compute architecture into a platform that compiler teams and data center operators can rely on. We shift complexity out of silicon and into software: the hardware abstraction layers, core system libraries, drivers, and runtime components that workloads enter the platform through. We build this stack in Rust. For system software living at the boundary between hardware and everything above it, we treat memory safety, explicit ownership, and long-lived API stability as the baseline rather than the goal - the foundation that lets us spend our judgment on the hard problems instead of on classes of bugs that should not exist.

As one of the principal engineers on this stack, set technical direction for the surfaces you own and shape the overall architecture alongside your fellow principals. Design the HAL, runtime interfaces, and data-movement pipelines the rest of the platform depends on; drive the hardest reliability and bring-up problems to root cause; and raise the throughput of the whole org by codifying the abstractions, patterns, and tooling that others build on. You will also help define how we engineer. We treat AI coding agents as a primary part of the workflow, and we expect our most senior engineers to be fluent in directing them - designing systems that are legible to both humans and agents, and turning hard-won judgment into leverage across the team.

What you’ll be doing:

  • Shape the architecture of the hardware abstraction layers and core system libraries, and own the API contracts for the components you lead.
  • Design and implement drivers, runtimes, and data movement and aggregation pipelines that execute workloads on novel silicon.
  • Build runtime interfaces for launching, monitoring, and managing workloads at production scale.
  • Drive triage of the most difficult sequencing, initialization, and cross-component runtime failures, and produce root-cause analyses that change how the system is built.
  • Lead new platform bring-up and NPI for new boards and silicon, in tight partnership with hardware engineering, compiler teams, and data center operations.
  • Multiply the team - establish the agent-assisted engineering practices, reusable abstractions, diagnostics, and documentation that let everyone move faster without destabilizing the platform.
  • Communicate architecture and design tradeoffs clearly, in writing and in diagrams, to audiences ranging from individual engineers to executive staff.

Requirements

  • MS in CS, CE, EE, or a related STEM field, or equivalent experience, and 12+ years building production system software.
  • Deep systems-programming expertise, with Rust as your language of choice for low-level work. You have shipped production Rust at the hardware or kernel boundary - drivers, firmware, runtimes, or similar - and you can articulate from experience where Rust earns its keep in system software and where it costs you. We work in Rust from day one; comfort is not enough, we want conviction.
  • A track record of designing and evolving libraries and APIs meant to be supported for years, including ABI and compatibility discipline.
  • Fluency in large, multi-repository codebases with layered dependencies.
  • Demonstrated leadership driving triage of difficult reliability issues to clear, written root-cause analysis.
  • Low-level platform experience: firmware and boot flows, RTOS, BMCs/MCUs, RISC-V, or closely related system software.
  • Linux driver or kernel-adjacent experience (for example, VFIO or similar subsystems).
  • Hardware bring-up and system triage experience: fault analysis, diagnostics, and validation in lab environments.
  • An established habit: building with AI coding agents - not as a novelty, but as a way you already ship and raise leverage. You can speak to how you design work to be agent-amenable and where you keep humans in the loop.

Ways to stand out from the crowd:

  • Experience having built Rust system software at the scale of a hyperscaler or a Rust-native hardware company - the kind of environment where Rust is the production language for low-level work, not an experiment.
  • Distributed systems experience: gRPC and RPC frameworks, coordination and telemetry patterns, MPI. Inference systems and token serving experience (vLLM or similar serving and runtime stacks) a huge plus.
  • Experience shipping and supporting customer-facing SDKs, including documentation and ABI compatibility practices.
  • Production readiness and delivery depth: CI/CD and release workflows, monitoring and alerting practices, Kubernetes, and data center operational workflows.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

About the company

Widely considered to be one of the technology world’s most desirable employers, NVIDIA has some of the most forward-thinking and hardworking people in the world inventing the future with us. Are you a creative and collaborative software lead seeking new challenges? If so, we want to hear from you! Join us and help build the real-time, cost-effective AI inference and computing platform that’s driving our success in this exciting and quickly growing field.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · World Congress 2021

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac · World Congress 2021

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

2:13 min

Evaluating Rust for modernizing embedded C firmware

Michael Friedrich Michael Friedrich · World Congress 2026 Europe

54 sec

Shifting focus to Rust for web backend development

Marco Otte-Witte · World Congress 2023

Videos

See all

Related articles

See all