Machine Learning Engineer (LLM Fine-Tuning & Inference) - FT - EMEA/APAC

Arc Full-time
United States
1 day ago
Apply on arc.dev
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Shift work
Job source

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Nvidia CUDA Continuous Integration Distributed Computing Environment Python (Programming Language) Linux System Administration Machine Learning Open Source Technology Performance Tuning QLoRA Graphics Processing Unit (GPU)
+17 more
Pytorch DeepSpeed Evaluation Pipelines Retrieval-Augmented Generation Transfer Learning Large Language Models Hallucination Detection Qwen Agentic-AI Fastapi Kubernetes HuggingFace Machine Learning Operations Mixed Precision Hardware Infrastructure Deepseek (AI Talent Sourcing Platform) Docker

Job description

We’re looking for a Machine Learning Engineer to fine-tune open-source LLMs and deploy them as fast, cost-efficient production services. You’ll own the full path from dataset preparation and training through quantisation, serving and performance optimisation., * Fine-tune open-source LLMs and SLMs (e.g., Llama, Qwen, Mistral, Gemma, DeepSeek) using SFT, LoRA/QLoRA and other PEFT techniques

  • Build, clean and validate domain-specific instruction datasets
  • Configure and run GPU training, including hyperparameters, mixed precision (BF16/FP16), gradient accumulation and checkpointing
  • Evaluate fine-tuned models against baselines using measurable criteria
  • Quantise and optimise models for production (AWQ/GPTQ, INT8/INT4, KV cache, continuous batching)
  • Deploy and serve self-hosted models with vLLM, TGI, Triton or similar
  • Benchmark and improve latency, throughput, GPU utilisation and inference cost
  • Expose models through production APIs (FastAPI or similar)

Requirements

PythonPyTorchHuggingfaceLlm inference tuning, * Strong Python and PyTorch

  • Hands-on experience fine-tuning LLMs with Hugging Face Transformers, PEFT and TRL (or Axolotl/Unsloth)
  • Solid grasp of transformer architecture and training fundamentals
  • Experience deploying LLMs on GPU infrastructure with vLLM, TGI, Triton or equivalent
  • Practical experience with quantisation and inference optimisation
  • Experience with Docker and Linux environments, * RAG, vector databases and AI agent development
  • Automated evaluation pipelines and LLM metrics (hallucination, faithfulness)
  • MLOps tooling: MLflow, W&B, Kubernetes, CI/CD for ML
  • AWS: SageMaker, EKS, Bedrock
  • Distributed training (DeepSpeed, multi-GPU), knowledge distillation, CUDA fundamentals
  • Experience with A100/H100/L40S GPUs

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on arc.dev
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:48 min

Evaluating methods for improving large language model outputs

Cedric Clyburn Cedric Clyburn +1 · World Congress 2024

1:59 min

Comparing Claude Sonnet and Qwen for text-to-SQL generation

Björn Buchhold Björn Buchhold · World Congress 2026 Europe

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

1:59 min

Building culturally aware LLMs for global audiences

Werner Vogels Werner Vogels +1 · World Congress 2026 Europe

2:50 min

Executing LoRA fine-tuning using serverless Databricks AI runtimes

Viktoria Semaan Viktoria Semaan · World Congress 2026 Europe

1:21 min

Running local coding agents on consumer hardware configurations

Chris Heilmann Chris Heilmann +2 · LIVE

Videos

See all

Related articles

See all