> Markdown version of [/jobs/ext/1916711-senior-solutions-architect-ai-factory-deployment-nvis](https://www.wearedevelopers.com/jobs/ext/1916711-senior-solutions-architect-ai-factory-deployment-nvis). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Solutions Architect, AI Factory Deployment - NVIS - **Company:** NVIDIA Ltd. - **Location:** Austin, TX, United States - **Experience:** Expert - **Salary:** $152,000.0 - $287,500.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Build Automation, Bash Shell, Computer Clusters, Software Debugging, Linux, Distributed Computing Environment, Distributed Systems, Python (Programming Language), Linux System Administration, Data Logging, Scripting, Large Language Models, Information Technology - **Published:** August 4, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=d65203b91d1022ad ## About the Role * Bachelor's degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or a related field. * 5+ years of experience managing Linux-based systems in HPC, distributed systems, or AI/ML environments. * Hands-on experience running AI/ML workloads on multi-GPU and/or multi-node clusters, including some exposure to NCCL. * Practical knowledge of collective communication patterns like AllReduce and AllToAll, and their application in ML/LLM training. * Skilled in Python and Shell/Bash for scripting, automation, and tooling. * Strong communication skills and the ability to work effectively with cross-functional teams. Ways to Stand Out From the Crowd: * Experience benchmarking distributed systems - crafting, running, and interpreting performance benchmarks. * Background in HPC performance engineering, SRE, or systems performance analysis for GPU-accelerated environments. * Familiarity with observability stacks (metrics/monitoring, logging, tracing) used for large distributed systems. * Experience building automation and CI-style pipelines for running and validating benchmarks at scale. * Demonstrated interest in using AI to solve practical problems, improve workflows, and guide data-driven decisions. ## Description We are in search of a curious and motivated Senior Solutions Architect to join our NVIDIA Infrastructure Specialists team. In this capacity, you'll support the creation, implementation, and verification of AI factories, focusing on running and debugging AI/LLM workloads and benchmarks on Linux-based GPU clusters. You'll engage with NCCL and collectives like AllReduce and AllToAll to boost performance and scalability, receiving mentorship from senior architects on the team. You will apply observability and automation to improve our benchmarking and validation efforts. You will be a key contact for troubleshooting workloads and benchmarks that fail, hang, or perform poorly. Additionally, you will collaborate with various NVIDIA teams to prepare AI factories for customers, validating both hardware and software for current AI applications. What You Will Be Doing: * Set up, adjust, and verify AI factory environments across multi-GPU and multi-node Linux clusters. * Validate configurations against guidelines for NCCL, collectives, and distributed training frameworks. * Run key AI/LLM benchmarks - setup, orchestration, result collection, and analysis. * Investigate and address problems when training jobs or benchmarks fail, hang, or perform below expectations. * Build and improve observability for AI factories (metrics, logs, traces, dashboards) to understand workload behavior and system health. * Build automation using Python and Shell for conducting benchmarks, retrieving results, and completing regression checks. * Analyze communication patterns and NCCL usage for AI/LLM workloads, concentrating on collectives such as AllReduce and AllToAll. * Help identify and recommend improvements to job configuration, parallelism strategies, and cluster settings to improve throughput, latency, and scaling efficiency. * Work closely with hardware, software, networking, and product teams to prepare AI factories for customer use. * Contribute to documentation and readiness materials for internal and customer-facing teams. ## Related Videos - [Your Next AI Needs 10,000 GPUs. Now What?](https://www.wearedevelopers.com/videos/1590-your-next-ai-needs-10-000-gpus-now-what) - [Old tools, new tricks](https://www.wearedevelopers.com/videos/1916-old-tools-new-tricks) - [JavaScript? No. Java Scripts! - Scripting with Java](https://www.wearedevelopers.com/videos/2094-javascript-no-java-scripts-scripting-with-java) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Building the Nervous System of AI - Michael Kagan (NVIDIA)](https://www.wearedevelopers.com/videos/2133-building-the-nervous-system-of-ai-michael-kagan-nvidia) - [MCP doesn’t suck — your agent does](https://www.wearedevelopers.com/videos/100202-mcp-doesn-t-suck-your-agent-does) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere)