Senior Solutions Architect, Physical AI Cloud

NVIDIA Ltd.
Santa Clara, CA, United States
1 day ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$152,000.0 - $241,500.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Airflow Amazon S3 Big Data Cloud Computing Cloud Engineering Computer Engineering Data Transformation Software Debugging DevOps Domain Name System (DNS)
+19 more
Software Engineering TCP/IP Workflow Management Systems Data Ingestion Spring Cloud Delivery Pipeline Large Language Models Backend Containerization AI Platforms Kubernetes Infrastructure Automation Frameworks Storage Technologies Information Technology Machine Learning Operations TensorRT Nim (Programming Language) Grpc Data Generation

Job description

We’re building a group of innovators to assist enterprises in deploying and accelerating NVIDIA’s three computer workloads for Physical AI. These include robotics simulation, synthetic data generation, multi-step model training, and inference, all on a large scale!

We are seeking a hands-on Solutions Architect with deep expertise in backend infrastructure, inference and cloud-native applications to design and scale Kubernetes-native environments for distributed Robotics workloads. This role offers an outstanding chance to build within the rapidly growing field of Robotics AI & Simulation. You’ll work closely with our product management, engineering, and business teams to drive the adoption of NVIDIA’s groundbreaking Physical AI technologies with our key ecosystem partners!

What you’ll be doing:

  • Help partners build scalable, observable, GPU-accelerated Physical AI pipelines through agentic workflows, cloud-native technologies, and NVIDIA frameworks such as OSMO.
  • Support development of Physical AI data factories for data ingestion, preprocessing, annotation, filtering, synthetic data generation, training, simulation, and evaluation.
  • Develop a deep understanding of robotics workload scaling and translate customer requirements into optimized cloud-native architectures, improving scheduling, cost, storage access, networking, and GPU utilization across hybrid infrastructure.
  • Accelerate distributed inference using NVIDIA technologies such as NIM, TensorRT-LLM, vLLM, and SGLang.
  • Collaborate with business, engineering, and product teams while providing technical guidance and mentorship to customers implementing Physical AI at scale.

Requirements

  • BS in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 5+ Years of experience in Solution Architecture or Infrastructure Engineering, advancing AI/ML systems from proof of concept to production on private/public cloud environments.
  • Experience with scaling Robotics workloads in one or more areas, such as multimodal model training, inference, robot learning and simulation, large scale data processing and generation.
  • Strong hands-on experience designing, deploying, and operating Kubernetes-based platforms for distributed GPU and AI workloads.
  • Expertise in networking (DNS, LB, TCP/IP, firewalls), storage technology, workflow orchestration softwares (Airflow, Argo, etc), modern DevOps practices (GitOps, IaC, Observability), and orchestrating efficient GPU workloads
  • Excellent communication skills to convey technical concepts to diverse audiences.

Ways to stand out from the crowd:

  • Hands-on experience with robotics frameworks (e.g., ROS2) and NVIDIA simulation and AI platforms such as Isaac Lab, Isaac Sim, GR00T or Cosmos.
  • Previous exposure to large scale Robotics data curation, annotation, filtering pipelines, including the use of AI models for data labeling.
  • Experience deploying NVIDIA inference technologies (Dynamo, NIM, Triton, vLLM) using acceleration techniques like quantization.
  • Proficiency using and developing agentic workflows to accelerate software development, infrastructure automation, troubleshooting, and deployment workflows.
  • Broad technical expertise across networking, compute, and storage systems (e.g., S3, NFS, Lustre), with hands-on experience building and debugging APIs (REST, gRPC).

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.jofdav.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · World Congress 2026 Europe

5:02 min

Mapping distributed compute paradigms to modern vehicles

Joachim Werner · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

4:56 min

Establishing internal service communication with gRPC

Florian Bader Florian Bader · World Congress 2026 Europe

3:50 min

Queues in TCP stacks and continuous network connections

Clemens Vasters Clemens Vasters · World Congress 2022

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all