LLM & Generative AI Engineer

Rolejoin
London, UK
10 days ago
Apply on www.apply4u.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Clean Code Principles Natural Language Processing Open Source Technology Search Technologies Software Construction Pytorch Large Language Models Deep Learning Generative AI Low Latency HuggingFace Machine Learning Operations

Requirements

About the RoleJoin our AI Engineering division in London to specialize in LLM fine-tuning, retrieval-augmented generation (RAG), and hosting private models. You will be responsible for tailoring deep learning models to specialized domain tasks.Key ResponsibilitiesFine-tune open-source models (Llama, Mistral, Qwen) for specific domain functionsOptimize model deployment pipelines for low latency and high throughputBuild advanced context management and semantic search solutionsImplement prompt evaluation frameworks and guardrail architecturesRequirements3+ years of experience focusing on Natural Language Processing and Generative AIHands-on experience with PyTorch, Hugging Face Transformers, and parameter-efficient fine-tuning (PEFT/LoRA)Experience deploying models with vLLM, Ollama, or Triton Inference ServerStrong background in software engineering best practices and clean code #J-18808-Ljbffr

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.apply4u.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:48 min

Balancing delivery latency with stream reliability and scale

Phil Cluff · LIVE

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

3:37 min

Accessing API documentation and testing remote driving latency

Alexandru Ciinaru Alexandru Ciinaru +3 · World Congress 2025

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

1:59 min

Building culturally aware LLMs for global audiences

Werner Vogels Werner Vogels +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all