Principal Solutions Architect

Rafay Systems, Inc.
United States
25 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Distributed Computing Environment Monitoring of Systems Identity and Access Management Python (Programming Language) Role-Based Access Control Tensorflow Prometheus
+12 more
Azure Machine Learning AI Infrastructure Pytorch Autoscaling Large Language Models Grafana Kubernetes Slurm Machine Learning Operations Hardware Infrastructure Data Pipelines Golang

Job description

Rafay seeks a Principal Solutions Architect to enable enterprise customers/Neo clouds in deploying, operating, and scaling AI/ML workloads on our GPU Platform-as-a-Service offering. This is a hybrid technical leadership and people-management role: alongside hands-on, customer-facing architecture work, this person will build, lead, and grow a team of Solutions Architects, collaborating with platform engineering, MLOps, data science, and infrastructure teams to architect production-ready AI infrastructure solutions built on Kubernetes and GPU-accelerated environments., Team Leadership & People Management

  • Recruit, hire, and onboard Solutions Architects as the team scales
  • Directly manage a team of Solutions Architects, including workload allocation, coaching, and day-to-day support
  • Set individual and team goals; conduct regular 1:1s and performance reviews
  • Own career development planning for direct reports, including skills growth, promotion readiness, and succession planning
  • Foster an inclusive, high-performing team culture aligned with Rafay’s values
  • Manage team capacity, prioritization, and staffing against customer and project demand
  • Partner with sales, engineering, and executive leadership on hiring plans and team structure
  • Mentor and upskill both direct reports and junior team members across the broader organization

Technical & Customer-Facing Responsibilities

  • Design comprehensive AI/ML platform architectures covering inference, training, and data pipelines
  • Develop reference architectures for GPU cluster deployment and LLM serving infrastructure
  • Evaluate inference serving frameworks including vLLM, TGI, and Triton
  • Advise on GPU fabric topology options for distributed training scenarios
  • Design observability strategies using DCGM, OpenTelemetry, and eBPF
  • Translate infrastructure requirements into actionable platform designs
  • Deliver technical presentations, workshops, and proof-of-concept engagements
  • Serve as trusted advisor on AI infrastructure strategy, cost optimization, and scaling
  • Partner with customer stakeholders to understand workload requirements
  • Architect networking, identity management, observability, and security integrations
  • Monitor and troubleshoot production environments for GPU utilization and cluster health
  • Lead root cause analysis for complex customer issues
  • Document reference architectures and implementation best practices

Requirements

  • 8+ years in infrastructure, platform, or solutions engineering roles
  • 3+ years focused on AI/ML infrastructure or MLOps
  • 2+ years of direct people-management experience, including hiring, performance management, and career development of technical staff
  • Demonstrated ability to lead and grow a technical team while remaining hands-on with customers and architecture
  • Deep Kubernetes expertise including cluster lifecycle and RBAC
  • Hands-on experience with NVIDIA GPU infrastructure (H100/H200/B200 preferred)
  • Proficiency with distributed training concepts (NCCL, tensor parallelism)
  • Experience with LLM inference serving and optimization
  • Familiarity with GPU Operator, MIG, SR-IOV, and network fabrics
  • Strong scripting and automation skills (Python, Bash, Go preferred)
  • Ability to communicate complex technical concepts to diverse audiences, including executive stakeholders
  • Experience with AWS, Azure, or GCP platforms
  • Familiarity with monitoring tools like Prometheus, Grafana, and OpenTelemetry
  • Understanding of GPU-based workloads and model serving
  • Proven troubleshooting capabilities for infrastructure issues
  • Excellent communication, coaching, and customer-facing skills, * Experience building a Solutions Architecture or technical pre-sales team from the ground up
  • Formal people-management training or leadership certification
  • Enterprise customer support experience in cloud-native environments
  • Familiarity with PyTorch and TensorFlow frameworks
  • Experience with Run:AI and Slurm
  • GPU scheduling and autoscaling expertise
  • Multi-tenant Kubernetes environment knowledge
  • MLOps platform experience
  • Technical workshop leadership experience
  • Relevant certifications (CKA, CKAD, AWS/Azure/GCP Solutions Architect)
  • Understanding of multi-tenant GPU isolation technologies

Benefits & conditions

Rafay is at the forefront of GPU PaaS technologies and Kubernetes and we offer unique opportunities to join a winning team working on foundational technology for cloud and AI/ML services and enterprises. We work in a collaborative environment that rewards creative thinking and provides opportunities to advance professional careers in advanced technology development. On top of this we offer a fun and dynamic work environment, a competitive salary, robust benefits and attractive stock options. As the first of our kind, we are truly in a class of our own.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Loading talks and stories from around this role…