> Markdown version of [/jobs/ext/1419421-senior-machine-learning-engineer-llm-inference-optimization](https://www.wearedevelopers.com/jobs/ext/1419421-senior-machine-learning-engineer-llm-inference-optimization). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Machine Learning Engineer, LLM Inference Optimization - **Company:** Nebius Inc. - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $195,200.0 - $262,200.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Nvidia CUDA, Software Debugging, Distributed Systems, Python (Programming Language), Machine Learning, Pytorch, Large Language Models, Low Latency, Free and Open-Source Software, TensorRT, Decoding - **Published:** July 24, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=8745a4fdc2dd4a42 ## About the Role * Strong Python and PyTorch engineering skills. * Hands-on experience deploying or optimizing LLM, VLM, or high-throughput transformer inference systems. * Practical knowledge of at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, or equivalent internal systems. * Strong understanding of transformer inference bottlenecks, including KV cache, attention, memory bandwidth, batching, parallelism, and long-context serving. * Ability to reason quantitatively about latency, throughput, quality, utilization, and cost tradeoffs. * Strong communication skills and ability to collaborate with research, kernel, infrastructure, product, and customer teams., * Experience with quantization-aware training, post-training quantization, FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ, SmoothQuant, or related techniques. * Experience with distillation, speculative decoding, EAGLE, Medusa, multi-token prediction, or other inference acceleration methods. * Experience with agentic workloads, including tool calling, structured outputs, streaming APIs, high concurrency, and multi-step orchestration. * CUDA or Triton familiarity, even if the role is not primarily a kernel-engineering role. * Open-source contributions to vLLM, SGLang, TensorRT-LLM, FlashInfer, LMCache, PyTorch, Triton, Ray, KServe, or related projects., Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. ## Description Nebius Token Factory is building an AI training and model post-training capability for frontier model improvement. This role owns the infrastructure that makes large-scale training and RL experiments possible, reliable, reproducible, and efficient. The work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering. A Senior MLE owns substantial model and endpoint optimization projects end to end. They are deeply hands-on, can debug difficult serving problems independently, and can deliver measurable improvements without needing heavy supervision. Your responsibilities: * Own optimization work for specific model families, customer endpoints, or serving backends. * Run engine comparisons and recommend practical serving configurations for specific workloads. * Debug model quality or performance regressions during production rollouts. * Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token. * Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems. * Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery. * Implement or integrate speculative decoding, draft-model approaches, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving. * Build reproducible benchmark harnesses for TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token. * Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers. * Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Decode Your People: Using PCM to Build High-Performance Teams](https://www.wearedevelopers.com/videos/100194-decode-your-people-using-pcm-to-build-high-performance-teams) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)