ML Infrastructure Engineer
Yobi & Tylo LLC
United States
about 2 months ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.indeed.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source
Tech stack
Large Language Models
Machine Learning Operations
TensorRT
Job description
We serve inference at $/token margins that don’t tolerate sloppy stacks. You’ll own the serving layer - vLLM, TensorRT-LLM, Triton - and the benchmarking discipline that keeps it honest., * Serving-stack selection per workload (continuous batching vs. static, KV cache strategy, paged attention).
- Quantization (FP8, AWQ, GPTQ) and the eval harness that proves the trade-offs.
- The InferenceBench-style benchmarks that compare our serving against the field.
Requirements
- Shipped at least one production inference stack on H100s or A100s.
- Read a recent vLLM PR for fun.
- Strong opinions about speculative decoding.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
almost 3 years ago
KD
Krissy Davis
The Best Large Language Models on The Market
almost 3 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
25 days ago
BB
Benedikt Bischof
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production
about 4 years ago
CH
Chris Heilmann
Dev Digest 132 - Binging WADFlix?
almost 2 years ago
DC
Daniel Cranney
Dev Digest 210: AI Agents Are Go! Is MCP Dead? LLMs Crack Anonymity
6 months ago