> Markdown version of [/jobs/ext/3570282-machine-learning-engineer-llm-fine-tuning-inference-ft-emea-apac](https://www.wearedevelopers.com/jobs/ext/3570282-machine-learning-engineer-llm-fine-tuning-inference-ft-emea-apac). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer (LLM Fine-Tuning & Inference) - FT - EMEA/APAC - **Company:** Arc Full-time - **Location:** United States (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Amazon Web Services, Nvidia CUDA, Continuous Integration, Distributed Computing Environment, Python (Programming Language), Linux System Administration, Machine Learning, Open Source Technology, Performance Tuning, QLoRA, Graphics Processing Unit (GPU), Pytorch, DeepSpeed, Evaluation Pipelines, Retrieval-Augmented Generation, Transfer Learning, Large Language Models, Hallucination Detection, Qwen, Agentic-AI, Fastapi, Kubernetes, HuggingFace, Machine Learning Operations, Mixed Precision, Hardware Infrastructure, Deepseek (AI Talent Sourcing Platform), Docker - **Published:** October 3, 2026 - **Apply:** https://arc.dev/signup?userType=developer&publicRandomKey=pomihw17k7&to=https%3A%2F%2Farc.dev%2Fdashboard%2Fd%2Fverified-jobs%2Fpomihw17k7%2Fapplication ## About the Role PythonPyTorchHuggingfaceLlm inference tuning, * Strong Python and PyTorch * Hands-on experience fine-tuning LLMs with Hugging Face Transformers, PEFT and TRL (or Axolotl/Unsloth) * Solid grasp of transformer architecture and training fundamentals * Experience deploying LLMs on GPU infrastructure with vLLM, TGI, Triton or equivalent * Practical experience with quantisation and inference optimisation * Experience with Docker and Linux environments, * RAG, vector databases and AI agent development * Automated evaluation pipelines and LLM metrics (hallucination, faithfulness) * MLOps tooling: MLflow, W&B, Kubernetes, CI/CD for ML * AWS: SageMaker, EKS, Bedrock * Distributed training (DeepSpeed, multi-GPU), knowledge distillation, CUDA fundamentals * Experience with A100/H100/L40S GPUs ## Description We're looking for a Machine Learning Engineer to fine-tune open-source LLMs and deploy them as fast, cost-efficient production services. You'll own the full path from dataset preparation and training through quantisation, serving and performance optimisation., * Fine-tune open-source LLMs and SLMs (e.g., Llama, Qwen, Mistral, Gemma, DeepSeek) using SFT, LoRA/QLoRA and other PEFT techniques * Build, clean and validate domain-specific instruction datasets * Configure and run GPU training, including hyperparameters, mixed precision (BF16/FP16), gradient accumulation and checkpointing * Evaluate fine-tuned models against baselines using measurable criteria * Quantise and optimise models for production (AWQ/GPTQ, INT8/INT4, KV cache, continuous batching) * Deploy and serve self-hosted models with vLLM, TGI, Triton or similar * Benchmark and improve latency, throughput, GPU utilisation and inference cost * Expose models through production APIs (FastAPI or similar) ## Related Videos - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [Theia AI Live Demo: Air Gapped AI for Developer Tools and IDEs](https://www.wearedevelopers.com/videos/100245-theia-ai-live-demo-air-gapped-ai-for-developer-tools-and-ides) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)