> Markdown version of [/jobs/ext/2062159-senior-machine-learning-engineer-llms](https://www.wearedevelopers.com/jobs/ext/2062159-senior-machine-learning-engineer-llms). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Machine Learning Engineer - LLMs - **Company:** Prosus - **Location:** Amsterdam, Netherlands - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Training Data, Big Data, Computer Clusters, Program Optimization, Code Review, Data Cleansing, Data Deduplication, Software Debugging, Distributed Computing Environment, Python (Programming Language), Machine Learning, Language Modeling, Node.Js, Software Deployment, Reinforcement Learning, Data Processing, Pytorch, Large Language Models, HuggingFace, TensorRT, Data Generation - **Published:** August 15, 2026 - **Apply:** https://www.adzuna.nl/details/5843190179 ## About the Role We're seeking a Senior Machine Learning Engineer to train domain-specific language models and provide technical leadership to the team. You'll own critical parts of our training infrastructure, mentor engineers, and drive technical decisions from data preparation through production deployment. You have deep hands-on experience training language models at scale, lead by example through rigorous experimentation and high-quality code, and are motivated by seeing your work deployed to millions of users. You thrive in fast-paced environments where you balance technical depth with practical business impact.What you'll do, * 7+ years of ML engineering experience * Technical leadership experience: mentoring engineers, conducting code reviews, making architecture decisions, and delivering projects with measurable business impact * Proven experience training and deploying language models to production (embedding models, encoder models, or large language models) including pre-training, continued pre-training, or fine-tuning with rigorous evaluation and inference optimization * Experience preparing large-scale training datasets: data filtering, quality assessment, deduplication strategies, and data mixture design * Hands-on experience with distributed training frameworks (DeepSpeed, FSDP, Megatron-LM, or Axolotl) including orchestrating multi-node jobs, debugging failures, and optimizing throughput * Strong understanding of training dynamics at scale: debugging loss instabilities, tuning learning rate schedules, managing training stability across long-running multi-node jobs * Expert Python and PyTorch with production experience using training libraries (Transformers, DeepSpeed, Accelerate), * Published research at ML conferences (NeurIPS, ICML, ICLR, ACL, EMNLP), released models on Hugging Face, created public benchmarks, or contributed to open-source projects * Experience with post-training methods: RLHF, DPO, GRPO, or other reinforcement learning approaches for alignment and instruction-following * Experience optimizing models for production inference including quantization, model compression, distillation, and serving frameworks (vLLM, TensorRT-LLM) * Understanding of memory optimization: gradient checkpointing, mixed precision training (FP16, BF16, FP8), ZeRO optimization * Deep knowledge of GPU architectures (A100, H100, H200) and their implications for training and inference optimization * Track record of building synthetic data generation pipelines for instruction tuning or domain adaptation ## Description * Analyze model performance and training data, formulate hypotheses, design and execute rigorous experiments to systematically improve model quality, training and inference efficiency, and downstream task performance * Drive technical decision-making for model architecture, training strategies, and infrastructure choices * Provide technical leadership and mentorship to ML engineers and interns, conducting code reviews, sharing best practices, and accelerating team growth * Train large language models through continued pre-training and full parameter fine-tuning on proprietary datasets * Build and optimize distributed training infrastructure across multi-node GPU clusters using frameworks like DeepSpeed, FSDP, Megatron-LM, or Axolotl * Own large-scale data preparation: filtering, quality assessment, deduplication, and data mixture strategies for training corpora at 100B+ token scale * Generate and curate high-quality synthetic data for instruction fine-tuning and capability enhancement * Debug training stability issues, optimize training and inference throughput (quantization, distillation, serving optimization), and monitor model performance throughout long-running distributed jobs * Build robust evaluation frameworks and establish metrics to measure model quality and guide decisions * Write production-grade, well-tested code and set engineering standards for the team ## Related Videos - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Adding knowledge to open-source LLMs](https://www.wearedevelopers.com/videos/1522-adding-knowledge-to-open-source-llms) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Inside the Mind of an LLM](https://www.wearedevelopers.com/videos/1617-inside-the-mind-of-an-llm) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)