Machine Learning Specialist

Stanford Black
Greater London, UK
5 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence C++ (Programming Language) Nvidia CUDA Distributed Computing Environment Distributed Systems Python (Programming Language) Machine Learning Performance Tuning Recommender Systems Tensorflow Software Engineering High Performance Computing
+9 more
Pytorch Large Language Models Deep Learning Parallel Computation Gpu Programming Kubernetes Information Technology Slurm Machine Learning Operations

Job description

  • We’re partnering with a highly quantitative research organisation building some of the most advanced machine learning systems in industry.
  • Engineers in this team operate at the intersection of machine learning, distributed systems, and high-performance computing, helping scale modern AI workloads across a large GPU estate. The work spans distributed training, inference optimisation, compute infrastructure, systems design, and performance engineering.
  • You’ll work directly with researchers to take cutting-edge ML ideas from prototype to production, solving problems that span software, hardware, networking, compilers, and large-scale distributed systems.
  • This is an opportunity to tackle technical challenges rarely seen outside leading AI labs and top-tier quantitative research firms.

Responsibilities

  • Design and optimise large-scale training and inference systems for modern ML workloads.
  • Improve throughput, latency, GPU utilisation and training efficiency across distributed environments.
  • Build infrastructure and tooling that accelerates experimentation and model development.
  • Partner with researchers to productionise novel ML approaches.
  • Drive performance improvements across software, hardware and networking layers.
  • Influence the technical direction of critical ML infrastructure used across the organisation.

Requirements

  • Strong experience in Machine Learning Engineering, Research Engineering, ML Infrastructure, Distributed Systems or Performance Engineering.
  • Excellent software engineering skills in Python and/or C++.
  • Experience working with modern ML frameworks such as PyTorch, JAX or TensorFlow.
  • Experience training, deploying or optimising large-scale machine learning models.
  • Strong understanding of distributed systems, parallel computing and performance optimisation.
  • Degree in Computer Science, Mathematics, Physics, Engineering or a related quantitative discipline, or equivalent industry experience.

Particularly Relevant Experience

  • Large-scale distributed training (DeepSpeed, FSDP, Megatron, Ray, DDP or similar).
  • GPU programming and optimisation (CUDA, Triton, NCCL, XLA, PTX).
  • Multi-GPU or multi-node training environments.
  • HPC, Kubernetes, Slurm or large-scale compute infrastructure.
  • Foundation models, LLMs, recommendation systems or large-scale deep learning.
  • Compiler technologies, kernel optimisation, inference optimisation or systems-level ML performance work.

Benefits & conditions

Why Join?

  • Work on some of the largest and most computationally intensive ML workloads in industry.
  • Solve challenging problems across distributed systems, GPU computing, machine learning infrastructure and performance optimisation.
  • Collaborate closely with exceptional researchers, engineers and quantitative scientists.
  • Significant autonomy and ownership from day one.
  • Deep investment in compute infrastructure and engineering excellence.
  • Competitive compensation and bonus structure.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · World Congress 2022

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all