ML Infrastructure Engineer

Yobi & Tylo LLC
United States
about 2 months ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Large Language Models Machine Learning Operations TensorRT

Job description

We serve inference at $/token margins that don’t tolerate sloppy stacks. You’ll own the serving layer - vLLM, TensorRT-LLM, Triton - and the benchmarking discipline that keeps it honest., * Serving-stack selection per workload (continuous batching vs. static, KV cache strategy, paged attention).

  • Quantization (FP8, AWQ, GPTQ) and the eval harness that proves the trade-offs.
  • The InferenceBench-style benchmarks that compare our serving against the field.

Requirements

  • Shipped at least one production inference stack on H100s or A100s.
  • Read a recent vLLM PR for fun.
  • Strong opinions about speculative decoding.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:08 min

Applying large language models to infrastructure tasks

Alfonso Sandoval Rosas Alfonso Sandoval Rosas · Europe 2026 Virtual

4:25 min

Optimizing agricultural practices through automated machine learning operations

Simi Olabisi · LIVE

3:32 min

Fundamentals and limitations of large language models

Krzystof Czieslak · LIVE

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · World Congress 2025

1:45 min

Introduction to serving large language models locally

Patrick Koss Patrick Koss · World Congress 2025

6:15 min

Executing open weight large language models with WebLLM

Christian Liebel Christian Liebel · World Congress 2025

Videos

See all

Related articles

See all