> Markdown version of [/jobs/ext/3625976-forward-deployed-engineer](https://www.wearedevelopers.com/jobs/ext/3625976-forward-deployed-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Forward Deployed Engineer - **Company:** Inferact, Inc. - **Location:** San Francisco, CA, United States - **Salary:** $200,000.0 - **Contract:** Permanent contract - **Skills:** Cloud Computing, Nvidia CUDA, Software Debugging, Programming Tools, Python (Programming Language), Open Source Technology, Software Engineering, TypeScript, AI Infrastructure, Graphics Processing Unit (GPU), Cloud Platform System, Large Language Models, Concurrency, Backend, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Low Latency, SGLang, Ray Serve, ROCm, Machine Learning Operations, TensorRT, Api Design, VLLM, Model Inference, Golang - **Published:** October 8, 2026 - **Apply:** https://startup.jobs/member-of-technical-staff-forward-deployed-engineer-inferact-10270395 ## About the Role * Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar. * Strong software engineering ability in Python, Go, TypeScript, or similar, with experience building production-quality integrations, tooling, services, automation, or prototypes. * Hands-on experience deploying or operating ML systems, model serving, AI infrastructure, cloud platforms, Kubernetes, or high-scale backend systems in production. * Ability to work directly with sophisticated customer engineering teams, understand ambiguous technical environments, and personally drive implementations and debugging to resolution. * Strong systems debugging skills across application, runtime, infrastructure, networking, identity, storage, observability, and distributed-system boundaries. * Ability to reason about latency, throughput, batching, model/runtime compatibility, scaling, reliability, and cost tradeoffs in production inference environments. * High ownership and strong technical communication, with the judgment to distinguish one-off customer work from problems that should become reusable product capabilities. Preferred qualifications: * Experience with vLLM, SGLang, TensorRT-LLM, TGI, Ray Serve, BentoML, or other LLM inference and model-serving systems. * Experience with NVIDIA or AMD GPUs, CUDA / ROCm, GPU scheduling, multi-GPU serving, or accelerator-backed infrastructure. * Experience deploying infrastructure software into enterprise, regulated, security-sensitive, or bring-your-own-cloud environments. * Experience building APIs, SDKs, CLIs, developer tooling, deployment platforms, control planes, or infrastructure products used by technical teams. * Experience profiling latency, throughput, concurrency, GPU utilization, bottlenecks, and performance regressions. ## Description We're looking for a Forward Deployed Engineer to make Inferact successful inside real customer environments. You'll work directly with customers to deploy, integrate, debug, and optimize vLLM-powered inference systems across cloud, Kubernetes, GPU, networking, and model-serving environments. This is a hands-on engineering role, not a traditional pre-sales position. You'll move from architecture discussions to implementation, own difficult production problems end-to-end, and work closely with core product and engineering teams to turn what you learn in the field into reusable product capabilities. Your work will directly affect customer time-to-value and how Inferact's platform evolves., * Contributed to open-source ML systems, inference infrastructure, cloud infrastructure, Kubernetes, or developer tooling. * Worked in a forward-deployed, customer engineering, field engineering, or highly technical solutions role where you personally wrote and shipped code. * Built deployment playbooks, reference architectures, automation, or tooling that materially reduced customer time-to-production. * Resolved severe customer-facing production issues that crossed multiple technical layers and required close partnership with core engineering. * Turned repeated customer problems into reusable product features, abstractions, documentation, or platform improvements.