> Markdown version of [/jobs/ext/1373370-ml-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/1373370-ml-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Infrastructure Engineer - **Company:** Yobi & Tylo LLC - **Location:** United States (Remote available) - **Contract:** Permanent contract - **Skills:** Large Language Models, Machine Learning Operations, TensorRT - **Published:** July 22, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=8935f6e8d4390550 ## About the Role * Shipped at least one production inference stack on H100s or A100s. * Read a recent vLLM PR for fun. * Strong opinions about speculative decoding. ## Description We serve inference at $/token margins that don't tolerate sloppy stacks. You'll own the serving layer - vLLM, TensorRT-LLM, Triton - and the benchmarking discipline that keeps it honest., * Serving-stack selection per workload (continuous batching vs. static, KV cache strategy, paged attention). * Quantization (FP8, AWQ, GPTQ) and the eval harness that proves the trade-offs. * The InferenceBench-style benchmarks that compare our serving against the field. ## Related Videos - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Creating Industry ready solutions with LLM Models](https://www.wearedevelopers.com/videos/899-creating-industry-ready-solutions-with-llm-models) - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) - [Unveiling the Magic: Scaling Large Language Models to Serve Millions](https://www.wearedevelopers.com/videos/1619-unveiling-the-magic-scaling-large-language-models-to-serve-millions) - [Lies, Damned Lies and Large Language Models](https://www.wearedevelopers.com/videos/1231-lies-damned-lies-and-large-language-models) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [Dev Digest 210: AI Agents Are Go! Is MCP Dead? LLMs Crack Anonymity](https://www.wearedevelopers.com/magazine/709-dev-digest-210-ai-agents-are-go-is-mcp-dead-llms-crack-anonymity)