Founding Engineer - ML Systems (up to £180k)

Dex
London, UK
7 days ago
Apply on www.apply4u.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
£180,000.0
Working hours
Regular working hours

Tech stack

Computer Clusters Distributed Computing Environment Python (Programming Language) Node.Js Workflow Management Systems Pytorch Large Language Models Backend Build Management Kubernetes Data Analytics Machine Learning Operations
+3 more
Api Design Data Pipelines Docker

Job description

This role is with one of Dex’s trusted partner companies. We work closely with their teams to truly understand their culture, goals, and what they’re looking for, so we can match you with the right opportunity and give you context about the role before you commit to a process. The role This company builds foundation models for extreme physics: the regimes behind semiconductors, aerospace, defence, and fusion energy. Their work accelerates progress in fields where existing simulation tools are too slow or brittle, enabling designs traditional workflows can’t reach. They’re backed by leading investors and tackling problems with global impact. You’ll be the first dedicated engineering owner for the ML stack, joining a small, highly technical team of researchers. This isn’t a narrow systems role, nor is it pure research; you’ll turn prototypes into robust code, integrate research branches, and build the in-house tooling for experiment tracking and hyperparameter optimisation. You’ll also add

Requirements

distributed training capabilities and own the backend platform delivering these models to customers, setting the engineering culture from day one. The work Own the entire ML stack: model code, training and evaluation workflows, and experiment infrastructure. Build and implement distributed training capabilities for large-scale model development. Integrate independently developed research branches into a coherent, production-ready codebase. Profile and resolve real bottlenecks in training stability and performance, improving system efficiency. Design and build the backend platform that delivers these foundation models to customers. What You Bring You are a senior, hands-on engineer who still writes and ships critical code. Strong practical experience with PyTorch across model code, data pipelines, and training loops. Proven track record with distributed training (multi-GPU/node, GPU clusters), understanding associated memory and communication challenges. Real model-engineering experience in physics/simulation, vision, or LLM systems, with a focus on data-driven system improvement. Solid Python platform and backend foundations: API design, workflow orchestration, and practical Docker/Kubernetes/IaC. #J-18808-Ljbffr

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.apply4u.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset · World Congress 2023

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

Videos

See all

Related articles

See all