Senior ML Solutions Architect - Token Factory

Jobgether
Veldhoven, Netherlands
1 day ago
Apply on www.adzuna.nl
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Data Cleansing DevOps High-Level Architecture Python (Programming Language) Machine Learning Open Source Technology Performance Tuning Azure Machine Learning AI Infrastructure
+12 more
Retrieval-Augmented Generation Large Language Models Prompt Engineering Model Validation Git AI Platforms Scikit Learn Kubernetes Deployment Automation TensorRT Serverless Computing Docker

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior ML Solutions Architect - Token Factory based in Netherlands.

Join a fast-growing AI infrastructure team building a serverless platform for running and customizing open-source LLMs in production. Help customers move from AI prototypes to scalable, reliable production applications without building their own complex inference stacks. Design optimized inference workflows and customized LLM solutions across multiple models and modalities. Work hands-on with inference, fine-tuning, evaluation, prompt engineering, and retrieval-augmented generation. Partner directly with customers to understand technical challenges and translate them into effective AI architectures. Collaborate closely with product and engineering teams to turn customer feedback into platform improvements. Work remotely across Europe in an international environment focused on high-impact AI projects, technical ownership, and continuous innovation. Accountabilities

  • Optimize LLM inference workflows across different modalities to deliver measurable business value and meet customer requirements.
  • Support customers with supervised and reinforcement-learning-based fine-tuning approaches to improve model quality and performance.
  • Design and implement LLM-powered solutions using serverless inference services and served open-source models.
  • Build production-ready applications using LLM APIs, including multimodal models covering text, vision, audio, and domain-specific use cases.
  • Provide technical guidance on prompt engineering, RAG architectures, model selection, inference optimization, and deployment strategies.
  • Guide customers through the transition from proof of concept to production, with a focus on performance, reliability, scalability, and cost efficiency.
  • Work closely with product and engineering teams to communicate customer needs, identify platform gaps, and contribute to roadmap development.
  • Help customers select appropriate models, inference configurations, and fine-tuning strategies based on their use cases and technical constraints.
  • Contribute to improving the platform and its capabilities by sharing practical insights from customer implementations and production workloads.

Requirements

  • 5+ years of professional experience working with ML/AI systems, including at least 2 years focused specifically on LLMs and generative AI.
  • Deep understanding of the modern LLM ecosystem, including model architectures, inference approaches, and fine-tuning techniques.
  • Hands-on experience running LLMs in production, including deploying and operating inference workloads at scale.
  • Strong practical experience with LLM fine-tuning, including supervised fine-tuning, SFT, LoRA, and data preparation or curation; experience with reinforcement-learning-based fine-tuning is a strong advantage.
  • Experience building LLM evaluation frameworks, including task-specific benchmarks, offline and online evaluation pipelines, and LLM-as-a-judge approaches.
  • Practical experience with modern inference frameworks and ML libraries such as vLLM, SGLang, TensorRT-LLM, or Transformers.
  • Experience deploying LLM-powered applications through APIs from providers such as OpenAI or Anthropic, as well as open-source models.
  • Strong Python programming skills and the ability to develop practical, production-oriented AI solutions.
  • Excellent communication skills, with the ability to explain complex technical concepts clearly to customers, engineers, product teams, and other audiences.
  • Experience working with multimodal AI models, such as vision-language or speech models, is a plus.
  • Familiarity with DevOps technologies including Docker, Kubernetes, and Git is beneficial.
  • Contributions to open-source ML or AI projects are an additional advantage.
  • Familiarity with cloud AI platforms such as AWS SageMaker or Bedrock, Google Vertex AI, or Azure ML is welcome.

Benefits & conditions

  • Competitive compensation.
  • Career growth and ongoing learning opportunities.
  • Flexible working arrangements with a high degree of autonomy and ownership.
  • Fully remote work opportunity from Europe.
  • Collaborative and innovative international working environment.
  • Opportunity to work on impactful AI infrastructure and production-grade LLM projects.
  • Exposure to advanced technologies across LLM inference, fine-tuning, evaluation, retrieval, and multimodal AI.
  • Opportunity to work closely with talented AI, engineering, product, and customer-facing teams.
  • Meaningful technical ownership and the opportunity to influence the evolution of an emerging AI platform.
  • Inclusive workplace committed to equal employment opportunities and a diverse working environment.
  • Reasonable accommodations are available during the application process where required.
  • Candidates must be authorized to work in the country where they apply and may need to provide proof of employment eligibility.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.nl
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:08 min

Applying large language models to infrastructure tasks

Alfonso Sandoval Rosas Alfonso Sandoval Rosas · Europe 2026 Virtual

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all