> Markdown version of [/jobs/ext/3294793-principal-solutions-architect](https://www.wearedevelopers.com/jobs/ext/3294793-principal-solutions-architect). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Solutions Architect - **Company:** Rafay Systems, Inc. - **Location:** United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Bash Shell, Cloud Computing, Distributed Computing Environment, Monitoring of Systems, Identity and Access Management, Python (Programming Language), Role-Based Access Control, Tensorflow, Prometheus, Azure Machine Learning, AI Infrastructure, Pytorch, Autoscaling, Large Language Models, Grafana, Kubernetes, Slurm, Machine Learning Operations, Hardware Infrastructure, Data Pipelines, Golang - **Published:** September 13, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=88bdfc7a1473883d ## About the Role * 8+ years in infrastructure, platform, or solutions engineering roles * 3+ years focused on AI/ML infrastructure or MLOps * 2+ years of direct people-management experience, including hiring, performance management, and career development of technical staff * Demonstrated ability to lead and grow a technical team while remaining hands-on with customers and architecture * Deep Kubernetes expertise including cluster lifecycle and RBAC * Hands-on experience with NVIDIA GPU infrastructure (H100/H200/B200 preferred) * Proficiency with distributed training concepts (NCCL, tensor parallelism) * Experience with LLM inference serving and optimization * Familiarity with GPU Operator, MIG, SR-IOV, and network fabrics * Strong scripting and automation skills (Python, Bash, Go preferred) * Ability to communicate complex technical concepts to diverse audiences, including executive stakeholders * Experience with AWS, Azure, or GCP platforms * Familiarity with monitoring tools like Prometheus, Grafana, and OpenTelemetry * Understanding of GPU-based workloads and model serving * Proven troubleshooting capabilities for infrastructure issues * Excellent communication, coaching, and customer-facing skills, * Experience building a Solutions Architecture or technical pre-sales team from the ground up * Formal people-management training or leadership certification * Enterprise customer support experience in cloud-native environments * Familiarity with PyTorch and TensorFlow frameworks * Experience with Run:AI and Slurm * GPU scheduling and autoscaling expertise * Multi-tenant Kubernetes environment knowledge * MLOps platform experience * Technical workshop leadership experience * Relevant certifications (CKA, CKAD, AWS/Azure/GCP Solutions Architect) * Understanding of multi-tenant GPU isolation technologies ## Description Rafay seeks a Principal Solutions Architect to enable enterprise customers/Neo clouds in deploying, operating, and scaling AI/ML workloads on our GPU Platform-as-a-Service offering. This is a hybrid technical leadership and people-management role: alongside hands-on, customer-facing architecture work, this person will build, lead, and grow a team of Solutions Architects, collaborating with platform engineering, MLOps, data science, and infrastructure teams to architect production-ready AI infrastructure solutions built on Kubernetes and GPU-accelerated environments., Team Leadership & People Management * Recruit, hire, and onboard Solutions Architects as the team scales * Directly manage a team of Solutions Architects, including workload allocation, coaching, and day-to-day support * Set individual and team goals; conduct regular 1:1s and performance reviews * Own career development planning for direct reports, including skills growth, promotion readiness, and succession planning * Foster an inclusive, high-performing team culture aligned with Rafay's values * Manage team capacity, prioritization, and staffing against customer and project demand * Partner with sales, engineering, and executive leadership on hiring plans and team structure * Mentor and upskill both direct reports and junior team members across the broader organization Technical & Customer-Facing Responsibilities * Design comprehensive AI/ML platform architectures covering inference, training, and data pipelines * Develop reference architectures for GPU cluster deployment and LLM serving infrastructure * Evaluate inference serving frameworks including vLLM, TGI, and Triton * Advise on GPU fabric topology options for distributed training scenarios * Design observability strategies using DCGM, OpenTelemetry, and eBPF * Translate infrastructure requirements into actionable platform designs * Deliver technical presentations, workshops, and proof-of-concept engagements * Serve as trusted advisor on AI infrastructure strategy, cost optimization, and scaling * Partner with customer stakeholders to understand workload requirements * Architect networking, identity management, observability, and security integrations * Monitor and troubleshoot production environments for GPU utilization and cluster health * Lead root cause analysis for complex customer issues * Document reference architectures and implementation best practices