Sr. Solutions Engineer

RAFAY CORPORATION
United States
2 days ago
Apply on www.thejobnetwork.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Distributed Systems Linux System Administration Networking Basics Prometheus Virtualization Technology Delivery Pipeline Grafana Containerization Kubernetes Infrastructure Automation Frameworks Information Technology
+5 more
Slurm Hardware Infrastructure VLLM Model Inference Terraform

Requirements

Rafay seeks a customer-focused, technically skilled Senior Solutions Engineer.. The role involves partnering with enterprise customers/Neo clouds to architect, deploy, and optimize cloud-native infrastructure using the Rafay platform and Kubernetes ecosystems, including GPU-accelerated environments for AI/ML and inference workloads., * 8+ years of experience in customer-facing technical roles, including implementation or consulting \n

  • Hands-on AI/ML engineering experience with model inference and deployment workflows \n

  • Hands-on experience with NVIDIA GPU infrastructure (H100/H200/B200 preferred) \n

  • Familiarity with GPU Operator, MIG, SR-IOV, and GPU fabric/networking fundamentals \n

  • Experience with LLM inference serving frameworks such as vLLM, TGI, or Triton \n

  • Strong Kubernetes and cloud-native environment expertise \n

  • Excellent written and verbal communication abilities \n

  • Deep troubleshooting skills across networking, virtualization, container orchestration, and public cloud platforms \n

  • Infrastructure automation and IaC (Terraform preferred) experience desired \n

  • Linux systems administration and distributed systems knowledge \n

  • Proven ability to lead technical projects and manage multiple engagements \n

  • Bachelor’s degree in Computer Science, Engineering, IT, or equivalent experience \n, * Experience with Run:AI and Slurm for GPU scheduling \n

  • GPU autoscaling and multi-tenant GPU isolation expertise

Benefits & conditions

  • Deploy and configure GPU-accelerated Kubernetes clusters for AI/ML workloads, including GPU Operator, MIG, and SR-IOV setup \n

  • Implement and validate LLM inference serving stacks (vLLM, TGI, Triton) during customer implementations \n

  • Troubleshoot GPU cluster health, utilization, and GPU fabric/networking issues (NCCL, RoCEv2) \n

  • Configure GPU observability using DCGM, OpenTelemetry, and related monitoring tooling \n

  • Advise customers on distributed training and inference topology options for NVIDIA GPU infrastructure (H100/H200/B200) \n

\n

\n

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.thejobnetwork.com
Prepare application

Good distractions

Loading talks and stories from around this role…