Machine Learning Specialist
Stanford Black
Greater London, UK
5 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.collegerecruiter.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source
Tech stack
Artificial Intelligence
C++ (Programming Language)
Nvidia CUDA
Distributed Computing Environment
Distributed Systems
Python (Programming Language)
Machine Learning
Performance Tuning
Recommender Systems
Tensorflow
Software Engineering
High Performance Computing
+9 more
Pytorch
Large Language Models
Deep Learning
Parallel Computation
Gpu Programming
Kubernetes
Information Technology
Slurm
Machine Learning Operations
Job description
- We’re partnering with a highly quantitative research organisation building some of the most advanced machine learning systems in industry.
- Engineers in this team operate at the intersection of machine learning, distributed systems, and high-performance computing, helping scale modern AI workloads across a large GPU estate. The work spans distributed training, inference optimisation, compute infrastructure, systems design, and performance engineering.
- You’ll work directly with researchers to take cutting-edge ML ideas from prototype to production, solving problems that span software, hardware, networking, compilers, and large-scale distributed systems.
- This is an opportunity to tackle technical challenges rarely seen outside leading AI labs and top-tier quantitative research firms.
Responsibilities
- Design and optimise large-scale training and inference systems for modern ML workloads.
- Improve throughput, latency, GPU utilisation and training efficiency across distributed environments.
- Build infrastructure and tooling that accelerates experimentation and model development.
- Partner with researchers to productionise novel ML approaches.
- Drive performance improvements across software, hardware and networking layers.
- Influence the technical direction of critical ML infrastructure used across the organisation.
Requirements
- Strong experience in Machine Learning Engineering, Research Engineering, ML Infrastructure, Distributed Systems or Performance Engineering.
- Excellent software engineering skills in Python and/or C++.
- Experience working with modern ML frameworks such as PyTorch, JAX or TensorFlow.
- Experience training, deploying or optimising large-scale machine learning models.
- Strong understanding of distributed systems, parallel computing and performance optimisation.
- Degree in Computer Science, Mathematics, Physics, Engineering or a related quantitative discipline, or equivalent industry experience.
Particularly Relevant Experience
- Large-scale distributed training (DeepSpeed, FSDP, Megatron, Ray, DDP or similar).
- GPU programming and optimisation (CUDA, Triton, NCCL, XLA, PTX).
- Multi-GPU or multi-node training environments.
- HPC, Kubernetes, Slurm or large-scale compute infrastructure.
- Foundation models, LLMs, recommendation systems or large-scale deep learning.
- Compiler technologies, kernel optimisation, inference optimisation or systems-level ML performance work.
Benefits & conditions
Why Join?
- Work on some of the largest and most computationally intensive ML workloads in industry.
- Solve challenging problems across distributed systems, GPU computing, machine learning infrastructure and performance optimisation.
- Collaborate closely with exceptional researchers, engineers and quantitative scientists.
- Significant autonomy and ownership from day one.
- Deep investment in compute infrastructure and engineering excellence.
- Competitive compensation and bonus structure.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.collegerecruiter.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
almost 3 years ago
BB
Benedikt Bischof
MLOps And AI Driven Development
over 4 years ago
BB
Benedikt Bischof
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production
about 4 years ago
LM
Luis Minvielle
How to Become an AI Engineer
almost 3 years ago
BB
Benedikt Bischof
MLOps – What’s the deal behind it?
almost 4 years ago
ER
Erin Rifkin
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
about 1 year ago