> Markdown version of [/jobs/ext/2304959-machine-learning-specialist](https://www.wearedevelopers.com/jobs/ext/2304959-machine-learning-specialist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Specialist - **Company:** Stanford Black - **Location:** Greater London, UK - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, C++ (Programming Language), Nvidia CUDA, Distributed Computing Environment, Distributed Systems, Python (Programming Language), Machine Learning, Performance Tuning, Recommender Systems, Tensorflow, Software Engineering, High Performance Computing, Pytorch, Large Language Models, Deep Learning, Parallel Computation, Gpu Programming, Kubernetes, Information Technology, Slurm, Machine Learning Operations - **Published:** August 30, 2026 - **Apply:** https://www.collegerecruiter.com/job/2815617112-machine-learning-specialist ## About the Role * Strong experience in Machine Learning Engineering, Research Engineering, ML Infrastructure, Distributed Systems or Performance Engineering. * Excellent software engineering skills in Python and/or C++. * Experience working with modern ML frameworks such as PyTorch, JAX or TensorFlow. * Experience training, deploying or optimising large-scale machine learning models. * Strong understanding of distributed systems, parallel computing and performance optimisation. * Degree in Computer Science, Mathematics, Physics, Engineering or a related quantitative discipline, or equivalent industry experience. Particularly Relevant Experience * Large-scale distributed training (DeepSpeed, FSDP, Megatron, Ray, DDP or similar). * GPU programming and optimisation (CUDA, Triton, NCCL, XLA, PTX). * Multi-GPU or multi-node training environments. * HPC, Kubernetes, Slurm or large-scale compute infrastructure. * Foundation models, LLMs, recommendation systems or large-scale deep learning. * Compiler technologies, kernel optimisation, inference optimisation or systems-level ML performance work. ## Description * We're partnering with a highly quantitative research organisation building some of the most advanced machine learning systems in industry. * Engineers in this team operate at the intersection of machine learning, distributed systems, and high-performance computing, helping scale modern AI workloads across a large GPU estate. The work spans distributed training, inference optimisation, compute infrastructure, systems design, and performance engineering. * You'll work directly with researchers to take cutting-edge ML ideas from prototype to production, solving problems that span software, hardware, networking, compilers, and large-scale distributed systems. * This is an opportunity to tackle technical challenges rarely seen outside leading AI labs and top-tier quantitative research firms. Responsibilities * Design and optimise large-scale training and inference systems for modern ML workloads. * Improve throughput, latency, GPU utilisation and training efficiency across distributed environments. * Build infrastructure and tooling that accelerates experimentation and model development. * Partner with researchers to productionise novel ML approaches. * Drive performance improvements across software, hardware and networking layers. * Influence the technical direction of critical ML infrastructure used across the organisation. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Getting Started with Machine Learning](https://www.wearedevelopers.com/videos/260-getting-started-with-machine-learning) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)