> Markdown version of [/jobs/ext/1523252-senior-software-developer-models-team-token-factory](https://www.wearedevelopers.com/jobs/ext/1523252-senior-software-developer-models-team-token-factory). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Developer: Models Team (Token Factory) - **Company:** Nebius - **Location:** Amsterdam, Netherlands - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Computer Programming, InfiniBand, Python (Programming Language), Autoscaling, Large Language Models, Kubernetes, Decoding - **Published:** July 15, 2026 - **Apply:** https://www.adzuna.nl/details/5778448052 ## About the Role * Experience serving LLMs in production * Strong Python and/or Go programming skills * Experience designing and operating highly scalable, highly available distributed services Nice to have: * Contributions to vLLM, SGLang, TRT-LLM, or NVIDIA ecosystem open-source projects * Deep understanding of KV cache management, speculative decoding, and quantization * Experience with LLM evaluation frameworks * Hands-on experience with performance benchmarking and optimization * Deep understanding of Kubernetes * Familiarity with distributed serving architectures and autoscaling * Knowledge of InfiniBand, RoCE, or high-performance networking, Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Unveiling the Magic: Scaling Large Language Models to Serve Millions](https://www.wearedevelopers.com/videos/1619-unveiling-the-magic-scaling-large-language-models-to-serve-millions) - [The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)](https://www.wearedevelopers.com/videos/2067-the-gashlycrumb-tinies-of-ai-networking-you-must-know-or-languish) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)