Sr. Solutions Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+5 more
Requirements
Rafay seeks a customer-focused, technically skilled Senior Solutions Engineer.. The role involves partnering with enterprise customers/Neo clouds to architect, deploy, and optimize cloud-native infrastructure using the Rafay platform and Kubernetes ecosystems, including GPU-accelerated environments for AI/ML and inference workloads., * 8+ years of experience in customer-facing technical roles, including implementation or consulting \n
-
Hands-on AI/ML engineering experience with model inference and deployment workflows \n
-
Hands-on experience with NVIDIA GPU infrastructure (H100/H200/B200 preferred) \n
-
Familiarity with GPU Operator, MIG, SR-IOV, and GPU fabric/networking fundamentals \n
-
Experience with LLM inference serving frameworks such as vLLM, TGI, or Triton \n
-
Strong Kubernetes and cloud-native environment expertise \n
-
Excellent written and verbal communication abilities \n
-
Deep troubleshooting skills across networking, virtualization, container orchestration, and public cloud platforms \n
-
Infrastructure automation and IaC (Terraform preferred) experience desired \n
-
Linux systems administration and distributed systems knowledge \n
-
Proven ability to lead technical projects and manage multiple engagements \n
-
Bachelor’s degree in Computer Science, Engineering, IT, or equivalent experience \n, * Experience with Run:AI and Slurm for GPU scheduling \n
-
GPU autoscaling and multi-tenant GPU isolation expertise
Benefits & conditions
-
Deploy and configure GPU-accelerated Kubernetes clusters for AI/ML workloads, including GPU Operator, MIG, and SR-IOV setup \n
-
Implement and validate LLM inference serving stacks (vLLM, TGI, Triton) during customer implementations \n
-
Troubleshoot GPU cluster health, utilization, and GPU fabric/networking issues (NCCL, RoCEv2) \n
-
Configure GPU observability using DCGM, OpenTelemetry, and related monitoring tooling \n
-
Advise customers on distributed training and inference topology options for NVIDIA GPU infrastructure (H100/H200/B200) \n
\n
\n
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Loading talks and stories from around this role…