Forward Deployment Engineer

GLINT TECH SOLUTIONS LLC
Mountain View, CA, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
$100,000.0 - $200,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Profiling Software Debugging Distributed Systems Python (Programming Language) Node.Js Azure Machine Learning Software Engineering Reinforcement Learning Graphics Processing Unit (GPU) Pytorch Large Language Models
+2 more
Kubernetes Hardware Infrastructure

Job description

We’re looking for a Forward Deployment Engineer (FDE) to work directly with customers and partners to design, deploy, and validate inference and reinforcement learning (RL) proof-of-concepts on GMI’s GPU infrastructure., Own customer POCs end-to-end

  • Deploy and optimize LLM inference, RL training, and post-training workflows on GMI clusters
  • Translate customer requirements into concrete system designs and experiments

Forward-deploy with customers

  • Work hands-on with research teams, startups, and enterprise customers
  • Debug performance, stability, and correctness issues in real environments

Inference deployment

  • Stand up and tune inference stacks (e.g. vLLM / SGLang / Ray Serve-style architectures)
  • Optimize latency, throughput, GPU utilization, and cost efficiency

RL & post-training POCs

  • Support RLHF / RFT / SFT workflows using customer-provided datasets
  • Integrate SDKs, training APIs, and cluster resources to shorten idea experiment cycles

Performance & reliability

  • Diagnose GPU, networking, and distributed system bottlenecks
  • Run benchmarks, profiling, and stress tests on multi-GPU / multi-node setups

Feedback loop to product

  • Feed real-world customer learnings back into GMI’s platform, SDKs, and APIs
  • Help shape reference architectures, cookbooks, and best practices

Requirements

Do you have experience in Technical troubleshooting support?, * Strong software engineering background (Python required; Go / Rust a plus)

  • Hands-on experience with ML inference or training systems
  • Familiarity with distributed systems and GPUs (multi-GPU, multi-node)
  • Comfort working directly with customers and ambiguous requirements
  • Ability to debug end-to-end systems (code, infra, networking, performance)

Nice to Have

  • Experience with:
  • LLM inference frameworks (vLLM, SGLang, Ray Serve, Triton, etc.)
  • RL or post-training workflows (RLHF, RFT, SFT)
  • PyTorch, DeepSpeed, Megatron-LM, or similar
  • Kubernetes-based ML platforms
  • GPU performance profiling and optimization
  • Prior experience as:
  • Forward Deployed Engineer
  • Solutions Engineer
  • ML Platform Engineer
  • Applied Research Engineer, * Engineers who like shipping over theorizing
  • People who enjoy being the last mile problem solver
  • Builders who want exposure to both deep systems and applied ML
  • Those excited by early-stage POCs that turn into real production systems

Benefits & conditions

$100,000 - $200,000 a year - Full-time

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset · WWC 2023

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

47 sec

Profiling native execution calls with async-profiler

Gonzalo Ortiz Jaureguizar Gonzalo Ortiz Jaureguizar · WWC Europe 2026

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:55 min

Identifying underlying Node.js runtime vulnerabilities using fuzzing tools

Sonya Moisset · WWC 2023

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all