> Markdown version of [/jobs/ext/1885267-senior-solutions-architect-generative-ai](https://www.wearedevelopers.com/jobs/ext/1885267-senior-solutions-architect-generative-ai). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Solutions Architect, Generative AI - **Company:** NVIDIA Ltd. - **Location:** Santa Clara, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $184,000.0 - $287,500.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Systems Engineering, Computer Clusters, Profiling, Computer Engineering, Network Congestion, Software Debugging, Linux, Microprocessors, Distributed Computing Environment, Distributed Systems, Network Topologies, InfiniBand, Python (Programming Language), Routing, Remote Direct Memory Access, Reliability Engineering, Runbook, Shell Script, AI Infrastructure, Graphics Processing Unit (GPU), Cloud Platform System, High Performance Computing, Generative AI, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Slurm, Data Pipelines - **Published:** August 1, 2026 - **Apply:** https://www.disabledperson.com/jobs/73949080-senior-solutions-architect-generative-ai ## About the Role * BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or another Engineering field, or equivalent experience. * 6+ years of experience in AI infrastructure, systems engineering, high-performance computing, networking, site reliability engineering, or a related technical role. * Deep understanding of Linux systems, distributed computing, GPU architectures, and the hardware and software components of large-scale AI clusters. * Hands-on experience designing, deploying, operating, or troubleshooting high-performance GPU networks in on-premises or cloud environments using technologies such as InfiniBand, RoCE, or GPUDirect RDMA. * Experience debugging NCCL communication and distributed collective performance, including topology, transport, congestion, routing, and host-level configuration issues. * Experience profiling AI workloads and identifying performance bottlenecks across compute, networking, storage, and orchestration layers. * Experience with cluster schedulers and orchestration platforms such as Kubernetes and Slurm, along with containers and production monitoring systems. * Proficiency with Python, shell scripting, or similar languages for infrastructure automation, benchmarking, and systems troubleshooting., * Experience architecting and operating large-scale production GPU clusters for distributed training or inference. * Deep expertise with NVIDIA infrastructure technologies such as DGX/HGX systems, NVLink, NVSwitch, NCCL, InfiniBand, and Spectrum-X. * Hands-on experience using tools and telemetry such as NCCL tests, DCGM, Nsight Systems, fabric counters, and host- or switch-level diagnostics to isolate performance and reliability issues. * Understanding of network topology, congestion control, collective communication patterns, and their impact on distributed AI workload performance. * Experience optimizing storage and data pipelines to sustain high-throughput training and inference workloads. ## Description NVIDIA is looking for an AI Solutions Architect with deep, hands-on experience in large-scale GPU systems. This role involves working with some of the world's leading consumer internet companies and frontier labs building foundation models. Primary responsibilities include accelerating customer workloads, designing high-performance AI infrastructure, and leading technical engagements around NVIDIA technologies. We work with the world's most successful technology companies, uniquely positioning you to observe and influence emerging infrastructure trends using the latest advancements. Join us in this exciting endeavor! What You'll Be Doing: * Collaborating closely with customers to maximize GPU utilization and end-to-end workload throughput while improving infrastructure reliability and reducing infrastructure costs. * Designing and optimizing large-scale AI clusters across GPU compute, high-performance networking, storage, workload scheduling, orchestration, and observability. * Profiling distributed training and inference workloads to identify bottlenecks across GPUs, CPUs, memory, network fabrics, storage systems, and software stack. * Diagnosing complex infrastructure and distributed systems issues spanning InfiniBand and RoCE fabrics, cloud interconnects, RDMA, NCCL, NVLink, and NVSwitch. * Leading proof-of-concepts and performance studies for large-scale AI infrastructure, developing benchmarking tools, automation, runbooks, and technical collateral as needed. * Partnering with NVIDIA's engineering, product, and sales teams to secure design wins and drive innovative solutions based on customer requirements and field feedback. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) ## Related Articles - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)