Manager, Solutions Architecture - GPU and Networking Systems

NVIDIA Corporation
Santa Clara, CA, United States
1 day ago
Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$224,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Systems Engineering C++ (Programming Language) Cloud Engineering Nvidia CUDA Computer Networks Computer Engineering Customer Data Management Data Centers Software Debugging Linux Ethernet
+12 more
Network Interface Controllers InfiniBand Linux Kernel Network Configuration and Change Management Software Architecture System Software Virtualization Technology Computer Networking Systems Computer Network Technologies Data Center Networking Information Technology Hardware Acceleration

Job description

NVIDIA is looking for a hands-on Solutions Architect Manager to lead a team of GPU, networking & software solution architects and engineers. Do you want to build and lead a group that designs, debugs, and deploys new AI hardware and software technologies into production in customer data centers? As part of the NVIDIA SA organization, you will drive people and technical leadership for end-to-end solutions deployments at some of NVIDIA’s most strategic technology customers, while directly contributing to designs and deep-dive debugging and shaping our product roadmap with customer feedback.

What you will be doing:

  • Recruit & manage a team of solutions architects, system/network and software engineers focused on large-scale GPU and AI networking deployments. Set priorities, allocate resources, mentor, and ensure high-quality customer delivery across multiple concurrent projects - while remaining directly involved in key technical reviews, design decisions, and critical debug efforts.
  • Provide deep subject-matter expertise in advanced GPU and network systems and serve as the senior technical point of contact for strategic customers. Personally lead and guide complex compute/network configuration and performance debugging, working side-by-side with your team to deliver performant, reliable clusters.
  • Guide your team as they lead network / compute / software architecture discussions, and support server, network, and cluster bring-up, including on-site data center work where needed.
  • Systematically collect and synthesize customer-specific requirements across your portfolio. Partner with GPU/Network Systems Engineering, Product Management, and Sales to influence roadmap priorities and packaging of reference designs and solutions.
  • Demonstrate SME in advanced GPU & network systems and be a trusted technical advisor to NVIDIA’s strategic customers. Bring customer-specific requirements to product teams to guide product roadmap features.
  • Identify new project opportunities for NVIDIA products and technology solutions in data center and AI applications. Work closely with the GPU/Network Systems Engineering, Product management and Sales teams

Requirements

  • BS/MS/PhD in Electrical/Computer Engineering, Computer Science, Physics, or other Engineering fields or equivalent experience.
  • 8+ overall years in Systems/Solutions/Field Engineering, Network or Data Center Engineering, or similar roles, with 2+ years leading or mentoring engineers or architects (formal manager or strong tech lead).
  • System-level expertise across CPU/GPU server architecture, NICs, Linux, system software, and kernel drivers. Experience with data center networking (Ethernet and/or InfiniBand switches, fabrics, and associated tooling). Familiarity with data center infrastructure (power, cooling, deployment constraints).
  • Proven ability to lead technical teams, set priorities, and drive complex projects from design through production. Demonstrated success working with Product Management, Sales, and Engineering.
  • Strong time management skills and ability to balance planning with hands-on support where needed.
  • Excellent written and verbal communication, including the ability to lead customer meetings, communicate status and risks, and produce clear design docs, debug summaries, and presentations.

Ways to stand out from the crowd:

  • Direct people management & recruiting experience for geographically distributed technical teams.
  • Track record leading bring-up and deployment of large clusters or supercomputing environments.
  • Background in external customer-facing roles (field engineering, escalations, or pre/post-sales architecture).
  • Systems engineering, coding, and debugging skills including experience with C/C++, Linux kernel and drivers
  • Hands-on experience with NVIDIA GPU systems and SDKs (e.g., CUDA), NVIDIA networking technologies (NICs, RoCE, InfiniBand), and/or ARM-based CPU solutions as well as familiarity with virtualization and cloud-native networking concepts.

We make extensive use of conferencing tools, but occasional (20%) travel is required for on-site visit to customers and industry events. We are open to remote work location and look forward to have you join our team!

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 3, and 272,000 USD - 431,250 USD for Level 4.

About the company

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you’re creative and autonomous, we want to hear from you!

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on nvidia.wd5.myworkdayjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

1:57 min

Routing cross-rack traffic seamlessly with NCCL

Kevin Klues Kevin Klues · World Congress 2025

4:52 min

Connecting namespaces with local virtual ethernet pairs

Oliver Seitz Oliver Seitz · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:51 min

Bypassing the CPU stack with remote direct memory access

Lerna Ekmekcioglu Lerna Ekmekcioglu · Europe 2026 Virtual

3:23 min

The AI workload technology stack and its components

Lerna Ekmekcioglu Lerna Ekmekcioglu · Europe 2026 Virtual

Videos

See all

Related articles

See all