Machine Learning Engineer (LLM Fine-Tuning & Inference) - FT - EMEA/APAC
Arc Full-time
United States
1 day ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on arc.dev
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Shift work
Job source
Tech stack
Application Programming Interfaces (APIs)
Amazon Web Services
Nvidia CUDA
Continuous Integration
Distributed Computing Environment
Python (Programming Language)
Linux System Administration
Machine Learning
Open Source Technology
Performance Tuning
QLoRA
Graphics Processing Unit (GPU)
+17 more
Pytorch
DeepSpeed
Evaluation Pipelines
Retrieval-Augmented Generation
Transfer Learning
Large Language Models
Hallucination Detection
Qwen
Agentic-AI
Fastapi
Kubernetes
HuggingFace
Machine Learning Operations
Mixed Precision
Hardware Infrastructure
Deepseek (AI Talent Sourcing Platform)
Docker
Job description
We’re looking for a Machine Learning Engineer to fine-tune open-source LLMs and deploy them as fast, cost-efficient production services. You’ll own the full path from dataset preparation and training through quantisation, serving and performance optimisation., * Fine-tune open-source LLMs and SLMs (e.g., Llama, Qwen, Mistral, Gemma, DeepSeek) using SFT, LoRA/QLoRA and other PEFT techniques
- Build, clean and validate domain-specific instruction datasets
- Configure and run GPU training, including hyperparameters, mixed precision (BF16/FP16), gradient accumulation and checkpointing
- Evaluate fine-tuned models against baselines using measurable criteria
- Quantise and optimise models for production (AWQ/GPTQ, INT8/INT4, KV cache, continuous batching)
- Deploy and serve self-hosted models with vLLM, TGI, Triton or similar
- Benchmark and improve latency, throughput, GPU utilisation and inference cost
- Expose models through production APIs (FastAPI or similar)
Requirements
PythonPyTorchHuggingfaceLlm inference tuning, * Strong Python and PyTorch
- Hands-on experience fine-tuning LLMs with Hugging Face Transformers, PEFT and TRL (or Axolotl/Unsloth)
- Solid grasp of transformer architecture and training fundamentals
- Experience deploying LLMs on GPU infrastructure with vLLM, TGI, Triton or equivalent
- Practical experience with quantisation and inference optimisation
- Experience with Docker and Linux environments, * RAG, vector databases and AI agent development
- Automated evaluation pipelines and LLM metrics (hallucination, faithfulness)
- MLOps tooling: MLflow, W&B, Kubernetes, CI/CD for ML
- AWS: SageMaker, EKS, Bedrock
- Distributed training (DeepSpeed, multi-GPU), knowledge distillation, CUDA fundamentals
- Experience with A100/H100/L40S GPUs
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on arc.dev
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
almost 3 years ago
BB
Benedikt Bischof
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production
about 4 years ago
KD
Krissy Davis
The Best Large Language Models on The Market
almost 3 years ago
LM
Luis Minvielle
How to Become an AI Engineer
almost 3 years ago
ER
Erin Rifkin
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
over 1 year ago
BB
Benedikt Bischof
MLOps And AI Driven Development
over 4 years ago