> Markdown version of [/jobs/ext/3391070-senior-solutions-engineer](https://www.wearedevelopers.com/jobs/ext/3391070-senior-solutions-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Solutions Engineer - **Company:** Rafay Systems, Inc. - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Distributed Computing Environment, Distributed Systems, Network Troubleshooting, Linux System Administration, Machine Learning, Networking Basics, Prometheus, Virtualization Technology, Delivery Pipeline, Large Language Models, Grafana, HybridCloud, Containerization, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Slurm, Machine Learning Operations, Hardware Infrastructure, Terraform - **Published:** September 13, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=a7e5428f267271a6 ## About the Role * 10+ years of experience in customer-facing technical roles, including implementation or consulting * Hands-on AI/ML engineering experience with model inference and deployment workflows * Hands-on experience with NVIDIA GPU infrastructure (H100/H200/B200 preferred) * Familiarity with GPU Operator, MIG, SR-IOV, and GPU fabric/networking fundamentals * Experience with LLM inference serving frameworks such as vLLM, TGI, or Triton * Strong Kubernetes and cloud-native environment expertise * Excellent written and verbal communication abilities * Deep troubleshooting skills across networking, virtualization, container orchestration, and public cloud platforms * Infrastructure automation and IaC (Terraform preferred) experience desired * Linux systems administration and distributed systems knowledge * Proven ability to lead technical projects and manage multiple engagements * Bachelor's degree in Computer Science, Engineering, IT, or equivalent experience, * Experience with Run:AI and Slurm for GPU scheduling * GPU autoscaling and multi-tenant GPU isolation expertise * Familiarity with DCGM-based GPU observability and Prometheus/Grafana dashboards * Relevant certifications (CKA, CKAD, or NVIDIA certifications) ## Description Rafay seeks a customer-focused, technically skilled Senior Solutions Engineer.. The role involves partnering with enterprise customers/Neo clouds to architect, deploy, and optimize cloud-native infrastructure using the Rafay platform and Kubernetes ecosystems, including GPU-accelerated environments for AI/ML and inference workloads., Core Implementation Responsibilities * Serve as primary technical lead for customer implementation engagements * Partner with customers on requirements gathering, architecture reviews, and solution design * Design and configure Rafay platform capabilities across public, private, and hybrid cloud environments * Troubleshoot complex infrastructure, networking, Kubernetes, and virtualization issues * Collaborate with Customer Success, Product, and Engineering teams on technical challenges * Manage customer issues within SLAs while maintaining expectations * Reproduce and analyze customer-reported issues; communicate findings to internal teams * Develop technical documentation, implementation guides, and runbooks * Mentor junior engineers and contribute to process improvements * Stay current on Rafay platform capabilities and emerging industry trends GPU & AI/ML Infrastructure * Deploy and configure GPU-accelerated Kubernetes clusters for AI/ML workloads, including GPU Operator, MIG, and SR-IOV setup * Implement and validate LLM inference serving stacks (vLLM, TGI, Triton) during customer implementations * Troubleshoot GPU cluster health, utilization, and GPU fabric/networking issues (NCCL, RoCEv2) * Configure GPU observability using DCGM, OpenTelemetry, and related monitoring tooling * Advise customers on distributed training and inference topology options for NVIDIA GPU infrastructure (H100/H200/B200)