Principal Solutions Architect
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+2 more
Requirements
-
10+ years in infrastructure, platform, or solutions engineering roles \n
-
3+ years focused on AI/ML infrastructure or MLOps \n
-
2+ years of direct people-management experience, including hiring, performance management, and career development of technical staff \n
-
Demonstrated ability to lead and grow a technical team while remaining hands-on with customers and architecture \n
-
Deep Kubernetes expertise including cluster lifecycle and RBAC \n
-
Hands-on experience with NVIDIA GPU infrastructure (H100/H200/B200 preferred) \n
-
Proficiency with distributed training concepts (NCCL, tensor parallelism) \n
-
Experience with LLM inference serving and optimization \n
-
Familiarity with GPU Operator, MIG, SR-IOV, and network fabrics \n
-
Strong scripting and automation skills (Python, Bash, Go preferred) \n
-
Ability to communicate complex technical concepts to diverse audiences, including executive stakeholders \n
-
Experience with AWS, Azure, or GCP platforms \n
-
Familiarity with monitoring tools like Prometheus, Grafana, and OpenTelemetry \n
-
Understanding of GPU-based workloads and model serving \n
-
Proven troubleshooting capabilities for infrastructure issues \n
-
Excellent communication, coaching, and customer-facing skills \n
Benefits & conditions
n \n
-
Recruit, hire, and onboard Solutions Architects as the team scales \n
-
Directly manage a team of Solutions Architects, including workload allocation, coaching, and day-to-day support \n
-
Set individual and team goals; conduct regular 1:1s and performance reviews \n
-
Own career development planning for direct reports, including skills growth, promotion readiness, and succession planning \n
-
Foster an inclusive, high-performing team culture aligned with Rafay’s values \n
-
Manage team capacity, prioritization, and staffing against customer and project demand \n
-
Partner with sales, engineering, and executive leadership on hiring plans and team structure \n
-
Mentor and upskill both direct reports and junior team members across the broader organization \n
\n
\n
Technical & Customer-Facing Responsibilities
\n \n
-
Design comprehensive AI/ML platform architectures covering inference, training, and data pipelines \n
-
Develop reference architectures for GPU cluster deployment and LLM serving infrastructure \n
-
Evaluate inference serving frameworks including vLLM, TGI, and Triton \n
-
Advise on GPU fabric topology options for distributed training scenarios \n
-
Design observability strategies using DCGM, OpenTelemetry, and eBPF \n
-
Translate infrastructure requirements into actionable platform designs \n
-
Deliver technical presentations, workshops, and proof-of-concept engagements \n
-
Serve as trusted advisor on AI infrastructure strategy, cost optimization, and scaling \n
-
Partner with customer stakeholders to understand workload requirements \n
-
Architect networking, identity management, observability, and security integrations \n
-
Monitor and troubleshoot production environments for GPU utilization and cluster health \n
-
Lead root cause analysis for complex customer issues \n
-
Document reference architectures and implementation best practices \n
\n
\n, n \n
-
Experience building a Solutions Architecture or technical pre-sales team from the ground up \n
-
Formal people-management training or leadership certification \n
-
Enterprise customer support experience in cloud-native environments \n
-
Familiarity with PyTorch and TensorFlow frameworks \n
-
Experience with Run:AI and Slurm \n
-
GPU scheduling and autoscaling expertise \n
-
Multi-tenant Kubernetes environment knowledge \n
-
MLOps platform experience \n
-
Technical workshop leadership experience \n
-
Relevant certifications (CKA, CKAD, AWS/Azure/GCP Solutions Architect) \n
-
Understanding of multi-tenant GPU isolation technologies \n
\n
\n
Why Join Rafay?
\n
\n
Rafay is at the forefront of GPU PaaS technologies and Kubernetes and we offer unique opportunities to join a winning team working on foundational technology for cloud and AI/ML services and enterprises. We work in a collaborative environment that rewards creative thinking and provides opportunities to advance professional careers in advanced technology development. On top of this we offer a fun and dynamic work environment, a competitive salary, robust benefits and attractive stock options. As the first of our kind, we are truly in a class of our own.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Loading talks and stories from around this role…