AI Infrastructure Engineer
Role details
Job location
Tech stack
Job description
You'll be a deeply technical specialist, who will architect and optimize distributed training across multiple GPUs and machines in AWS. You will eliminate bottlenecks in the data path to ensure training is fast and as capital efficient as possible alongside managing cluster orchestration using slurm and Kubernetes while preparing to expand into specialised GPU providers. And finally you will master the stack from pytorch based learning libraries to complex data networking and GPU utilisation.
Requirements
We're looking for an experienced AI Infrastructure or MLOps Engineer who has already built and operated production AI systems rather than someone at the beginning of their career. You'll have strong experience with the PyTorch ecosystem and a solid understanding of modern transformer architectures, including how to train and operate them at scale. Experience working with computer vision models and image-based datasets is particularly valuable, and we're looking for someone who thrives in an early-stage startup environment where ownership is high, pace is fast, and everyone contributes across the stack. This is a highly collaborative role, so you'll be excited to work alongside the team in the London- based lab every day.
Benefits & conditions
In return, you'll join at a pivotal stage where your work will directly shape both the technology and the future of the company. You'll be building cutting-edge robotics AI, while benefiting from a competitive salary, meaningful founding equity, and the opportunity to work alongside a world-class team in a state-of-the-art London lab.