Machine Learning Engineer Pytorch LLM

Client Server
London, UK
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Compensation
£110,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Computer Clusters Software Debugging Distributed Data Store Distributed Systems Memory Management Python (Programming Language) Machine Learning Open Source Technology Performance Tuning Software Engineering TypeScript
+8 more
Reinforcement Learning Pytorch Large Language Models Prompt Engineering Deep Learning Kubernetes Slurm Golang

Job description

Machine Learning Engineer (PyTorch LLM) London onsite to £110k

Do you have expertise with Machine Learning in production?

You could be progressing your career at a London based tech start-up with £5 million in recent pre-seed funding, in an impactful role that you’ll shape. The product is an AI agentic based platform that writes production grade code.

What’s in it for you: Salary to £110k Equity / stack options 30 days holiday (+ Bank Holidays) Daily lunch, monthly breakfasts Dog friendly office Pension Monthly socials Impactful role that you can shape and influence Your role:

As a Machine Learning Engineer you’ll take open-source LLMs (code and general models) and turn them into high-performance software engineer agents using supervised fine tuning and large scale reinforcement learning. This isn’t prompt engineering. You’ll design and run serious training experiments across multi-node GPU clusters, build RL loops where models write code and get rewarded (or penalised) by real test outcomes and push long-context and MoE style architectures to their limits.

You’ll work hands-on across the full stack: custom PyTorch dataloaders, distributed training (DDP/FSDP), experiment tracking, debugging NCCL issues at 2am, and squeezing performance out of multi-GPU jobs. You’ll help design opinionated reward functions that reflect what great engineering actually looks like, not just benchmark scores.

You’ll extend benchmark suites, test models on real world repositories, analyse failure modes and feed insights back into data and training strategy. Collaborating with infrastructure, product and research teams you’ll contribute to decisions about what to train next and how to measure results.

Location / WFH:

You’ll be based in the London, dog friendly office on a fulltime basis, with daily catered lunch, working hours (with no expectation to do more).

About you: You have strong experience with training deep learning models in production You have an indepth knowledge of PyTorch including hands-on experience with torch.distributed (DDP/FSDP-style training, distributed data loading, gradient scaling, etc.) You have experience of training large sequence models or LLMs at scale You have a software engineering background with Python, also familiar with TypeScript and / or Golang You have distributed systems / training ops experience including practical experience running multi-node jobs on GPU clusters (Slurm, Kubernetes, or managed cloud equivalents) and are familiar with GPU performance tuning: memory usage, mixed precision, throughput vs. latency tradeoffs You have experience within a reinforcement learning environment You’re collaborative with great communication skills You are degree educated to BSc / MSc in a relevant discipline Apply now or call to find out more about this Machine Learning Engineer opportunity.

Requirements

You have strong experience with training deep learning models in production You have an indepth knowledge of PyTorch including hands-on experience with torch.distributed (DDP/FSDP-style training, distributed data loading, gradient scaling, etc.) You have experience of training large sequence models or LLMs at scale You have a software engineering background with Python, also familiar with TypeScript and / or Golang You have distributed systems / training ops experience including practical experience running multi-node jobs on GPU clusters (Slurm, Kubernetes, or managed cloud equivalents) and are familiar with GPU performance tuning: memory usage, mixed precision, throughput vs. latency tradeoffs You have experience within a reinforcement learning environment You’re collaborative with great communication skills You are degree educated to BSc / MSc in a relevant discipline Apply now or call to find out more about this Machine Learning Engineer opportunity.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on tiptopjob.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

4:41 min

Replacing PyTorch with ONNX runtime for AWS Lambda deployments

Marek Suppa · LIVE

2:15 min

Open-source community and machine learning frameworks

Gian Marco Iodice Gian Marco Iodice · World Congress 2025

Videos

See all

Related articles

See all