Distinguished Engineer - Rack Scale Architecture

NVIDIA Corporation
Santa Clara, CA, United States
1 day ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$320,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Intelligent Platform Management Interface Computer Engineering Computer Graphics Data Centers Microprocessors Ethernet Firmware Field-Programmable Gate Array (FPGA) InfiniBand Node.Js Cloud Services
+7 more
Software Requirements Analysis Systems Architecture System Software Graphics Processing Unit (GPU) Cloud Platform System Computer Network Technologies Information Technology

Job description

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology-and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work.

NVIDIA has a rapidly expanding ecosystem of data center platform & node designs. From single node HGX/DGX systems all the way up to large multi-node NVLink domain rack architectures. These designs have become core to NVIDIA’s rapidly growing enterprise and cloud provider businesses. Each bringing together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. We’re searching for a highly motivated, technical leader to drive the engineering roadmap and innovation for our rack system software architecture. From firmware, kernel drivers, operating systems, networking, fabrics and associated user mode drivers + manageability software. You will work with component leads internally and engage with industry leading hyperscalar / cloud service providers on taking these products to market.

What you’ll be doing:

  • Drive the software end-to-end architecture for NVIDIA’s rack-scale products
  • Maintain deep understanding of the product portfolio and roadmap; translate forward-looking plans into clear, formal software requirements that anchor execution across the organization.
  • Ensure high quality & reliable software; serving as a trusted architectural partner to teams requiring guidance or oversight.
  • Work directly with major customers to understand their requirements and work to align their roadmap with NVIDIA’s roadmap.
  • Work with business partners and vendors to shape their products to meet NVIDIA’s needs.
  • Develop a roadmap of new technologies and protocols; drive their design and adoption.
  • Mentor architects and engineering teams to grow them into future leaders.
  • Make key technical decisions even when faced with ambiguity

Requirements

  • BS or MS degree in Computer Engineering, Computer Science, or related degree or equivalent experience.
  • 15+ years in the area of System architecture and design
  • Deep experience in designing architecture for scalable and performant server systems, particularly at the SW/HW interface.
  • Strong understanding of networking technology & protocols (e.g. Ethernet, Infiniband)
  • Previous experience working with complex system software for accelerators such as GPUs, DPUs, or FPGAs
  • Expertise in out-of-band and in-band management architectures.
  • Knowledge of system management protocols such as Redfish and IPMI.
  • Experience working with platform security experts to define tradeoffs between security and ease of use.
  • Demonstrable experience in implementing left shift strategy to de-risk program execution. Excellent written and verbal communication skills.

Ways to stand out from the crowd:

  • Knowledge of large-scale cloud and cluster level deployment and management systems. Experience with designing robust, resilient and performant scale-up fabrics
  • Demonstrated track record of leading data center products across the entire lifecycle, spanning inception, pre-silicon development, post-silicon bring-up, manufacturing, and deployment.
  • Familiarity with CXL, UCIE and other C2C technology architectures. Knowledge in storage and networking technologies.

We are widely considered to be one of the technology world’s most desirable employers, and as a result have some of the most forward-thinking and hardworking people in the world working for us. So if you’re clever, creative, and driven, we’d love to have you join the team.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 320,000 USD - 488,750 USD.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

2:19 min

Orchestrating over-the-air firmware updates for vehicle modules

Denis Grahovac · World Congress 2021

4:52 min

Connecting namespaces with local virtual ethernet pairs

Oliver Seitz Oliver Seitz · World Congress 2025

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset · World Congress 2023

3:05 min

Acquiring Mellanox to build cohesive AI factories

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

2:20 min

Utilizing custom firmware for variable torque manipulation

Daniel Meilak Daniel Meilak +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all