> Markdown version of [/jobs/ext/3570279-senior-applied-research-engineer-machine-learning-remote-uk-eu](https://www.wearedevelopers.com/jobs/ext/3570279-senior-applied-research-engineer-machine-learning-remote-uk-eu). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Applied Research Engineer - Machine Learning | Remote (UK/EU) - **Company:** SGI, Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Profiling, Nvidia CUDA, Computer Programming, Software Debugging, Distributed Computing Environment, Memory Management, Python (Programming Language), Machine Learning, Graphics Processing Unit (GPU), Pytorch, Low Latency - **Published:** October 3, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/pou9cow0hg ## About the Role * Strong understanding of modern ML architectures and large-scale training pipelines. * Hands-on experience running distributed training jobs on multi-GPU systems. * Advanced profiling and debugging across CPU, GPU, memory usage, latency, and inter-GPU communication. * Strong programming skills in Python. * Experience with model scaling and parallelisation strategies, including tensor and pipeline parallelism. Nice to have * Familiarity with NCCL, MPI, and distributed communication primitives. * Knowledge of PyTorch and Triton internals. * Programming experience with C++ and CUDA. ## Description You will work on novel technical challenges in large-scale model development and contribute to technology that is changing how major organisations operate. This is an opportunity to join a category-defining company at an early stage and help shape its trajectory., * Profile end-to-end distributed training runs to identify bottlenecks across compute, GPU memory, and inter-GPU communication. * Influence architectural decisions to improve efficiency and reliability of large-scale training jobs, including developing Triton/CUDA kernels when needed. * Design and implement model scaling, parallelisation, and memory optimisation techniques for training workloads with very large context sizes. * Collaborate closely with ML Researchers to diagnose architectural inefficiencies, ensure new research ideas scale efficiently in practice, and share internal knowledge on optimisation. * Drive productionisation and serving of models from the research side, including improving inference efficiency via techniques such as quantisation.