> Markdown version of [/jobs/ext/3007000-senior-ml-solutions-architect-token-factory](https://www.wearedevelopers.com/jobs/ext/3007000-senior-ml-solutions-architect-token-factory). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior ML Solutions Architect - Token Factory - **Company:** Jobgether - **Location:** Veldhoven, Netherlands (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Data Cleansing, DevOps, High-Level Architecture, Python (Programming Language), Machine Learning, Open Source Technology, Performance Tuning, Azure Machine Learning, AI Infrastructure, Retrieval-Augmented Generation, Large Language Models, Prompt Engineering, Model Validation, Git, AI Platforms, Scikit Learn, Kubernetes, Deployment Automation, TensorRT, Serverless Computing, Docker - **Published:** September 20, 2026 - **Apply:** https://www.adzuna.nl/details/5890836249 ## About the Role * 5+ years of professional experience working with ML/AI systems, including at least 2 years focused specifically on LLMs and generative AI. * Deep understanding of the modern LLM ecosystem, including model architectures, inference approaches, and fine-tuning techniques. * Hands-on experience running LLMs in production, including deploying and operating inference workloads at scale. * Strong practical experience with LLM fine-tuning, including supervised fine-tuning, SFT, LoRA, and data preparation or curation; experience with reinforcement-learning-based fine-tuning is a strong advantage. * Experience building LLM evaluation frameworks, including task-specific benchmarks, offline and online evaluation pipelines, and LLM-as-a-judge approaches. * Practical experience with modern inference frameworks and ML libraries such as vLLM, SGLang, TensorRT-LLM, or Transformers. * Experience deploying LLM-powered applications through APIs from providers such as OpenAI or Anthropic, as well as open-source models. * Strong Python programming skills and the ability to develop practical, production-oriented AI solutions. * Excellent communication skills, with the ability to explain complex technical concepts clearly to customers, engineers, product teams, and other audiences. * Experience working with multimodal AI models, such as vision-language or speech models, is a plus. * Familiarity with DevOps technologies including Docker, Kubernetes, and Git is beneficial. * Contributions to open-source ML or AI projects are an additional advantage. * Familiarity with cloud AI platforms such as AWS SageMaker or Bedrock, Google Vertex AI, or Azure ML is welcome. ## Description This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior ML Solutions Architect - Token Factory based in Netherlands. Join a fast-growing AI infrastructure team building a serverless platform for running and customizing open-source LLMs in production. Help customers move from AI prototypes to scalable, reliable production applications without building their own complex inference stacks. Design optimized inference workflows and customized LLM solutions across multiple models and modalities. Work hands-on with inference, fine-tuning, evaluation, prompt engineering, and retrieval-augmented generation. Partner directly with customers to understand technical challenges and translate them into effective AI architectures. Collaborate closely with product and engineering teams to turn customer feedback into platform improvements. Work remotely across Europe in an international environment focused on high-impact AI projects, technical ownership, and continuous innovation. Accountabilities * Optimize LLM inference workflows across different modalities to deliver measurable business value and meet customer requirements. * Support customers with supervised and reinforcement-learning-based fine-tuning approaches to improve model quality and performance. * Design and implement LLM-powered solutions using serverless inference services and served open-source models. * Build production-ready applications using LLM APIs, including multimodal models covering text, vision, audio, and domain-specific use cases. * Provide technical guidance on prompt engineering, RAG architectures, model selection, inference optimization, and deployment strategies. * Guide customers through the transition from proof of concept to production, with a focus on performance, reliability, scalability, and cost efficiency. * Work closely with product and engineering teams to communicate customer needs, identify platform gaps, and contribute to roadmap development. * Help customers select appropriate models, inference configurations, and fine-tuning strategies based on their use cases and technical constraints. * Contribute to improving the platform and its capabilities by sharing practical insights from customer implementations and production workloads. ## Related Videos - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [How to Avoid LLM Pitfalls - Mete Atamel and Guillaume Laforge](https://www.wearedevelopers.com/videos/1328-how-to-avoid-llm-pitfalls-mete-atamel-and-guillaume-laforge) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market)