Solutions Architect, Inference Deployments

NVIDIA Ltd.
Santa Clara, CA, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$152,000.0 - $241,500.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Artificial Neural Networks Program Optimization DevOps Distributed Systems Dynamic Random-Access Memory Open Source Technology Remote Direct Memory Access Large Language Models Kubernetes TensorRT Nim (Programming Language)
+1 more
Decoding

Job description

We’re forming a team of innovators to roll out and enhance AI inference solutions at scale, demonstrating NVIDIA’s GPU technology and Kubernetes. As a Solutions Architect focused on inference, you’ll collaborate closely with our engineering, DevOps, and customers to develop enterprise AI solutions. Together, we’ll deliver generative AI to production!

What you’ll be doing:

  • Build inference pipelines with tools like NVIDIA Dynamo, distributing tasks among GPU workers to improve efficiency.

  • Collaborate with DevOps teams to orchestrate disaggregated inference using Kubernetes for complex workloads.

  • Accelerate inference pipelines using TensorRT-LLM, vLLM, SGLang, and other backends to ensure seamless integration with disaggregated inference.

  • Provide mentorship and technical leadership to customers and internal teams, guiding them through the deployment of disaggregated inference systems and resolving complex issues.

Requirements

  • 5+ Years in Solutions Architecture with a proven track record of deploying distributed systems and AI inference workloads on Kubernetes.

  • Experience with one of NVIDIA Dynamo, Triton Inference Server, or TensorRT-LLM for model optimization and serving.

  • GPU orchestration using NVIDIA GPU Operator, NIM Operator, and Multi-Instance GPU (MIG) partitioning.

  • Solving sophisticated GPU allocation, memory hierarchies (HBM, DRAM, SSD), and low-latency networking (RDMA, UCX).

  • Demonstrated success in tuning large language models for low-latency inference in enterprise environments.

  • BS in CS/Engineering or equivalent experience.

Ways to stand out from the crowd:

  • Prior experience deploying NVIDIA inference technologies such as Dynamo, NIM, NIXL and Grove.

  • Deep understanding of transformer neural network, and inference acceleration technologies like quantization, speculative decoding, WideEP etc.

  • NVIDIA Certified AI Engineer or similar credentials.

  • Contributions to open-source projects including NVIDIA Dynamo, vLLM, KServe, or SGLang.

Benefits & conditions

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on juju.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:33 min

Architecting CUDA and the AI software stack

Michael Kagan Michael Kagan +1 · WWC Europe 2026

2:03 min

Accelerating token generation speeds with speculative decoding techniques

Christin Pohl Christin Pohl · WWC Europe 2026

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:32 min

Core libraries driving inference engines and multi-GPU networking

Adolf Hohl Adolf Hohl · WWC 2024

3:10 min

Scaling customized inference models with NVIDIA NIM

Anshul Jindal Anshul Jindal · WWC 2025

Videos

See all

Related articles

See all