Founding Engineer - ML Systems (up to £180k)
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+3 more
Job description
This role is with one of Dex’s trusted partner companies. We work closely with their teams to truly understand their culture, goals, and what they’re looking for, so we can match you with the right opportunity and give you context about the role before you commit to a process. The role This company builds foundation models for extreme physics: the regimes behind semiconductors, aerospace, defence, and fusion energy. Their work accelerates progress in fields where existing simulation tools are too slow or brittle, enabling designs traditional workflows can’t reach. They’re backed by leading investors and tackling problems with global impact. You’ll be the first dedicated engineering owner for the ML stack, joining a small, highly technical team of researchers. This isn’t a narrow systems role, nor is it pure research; you’ll turn prototypes into robust code, integrate research branches, and build the in-house tooling for experiment tracking and hyperparameter optimisation. You’ll also add
Requirements
distributed training capabilities and own the backend platform delivering these models to customers, setting the engineering culture from day one. The work Own the entire ML stack: model code, training and evaluation workflows, and experiment infrastructure. Build and implement distributed training capabilities for large-scale model development. Integrate independently developed research branches into a coherent, production-ready codebase. Profile and resolve real bottlenecks in training stability and performance, improving system efficiency. Design and build the backend platform that delivers these foundation models to customers. What You Bring You are a senior, hands-on engineer who still writes and ships critical code. Strong practical experience with PyTorch across model code, data pipelines, and training loops. Proven track record with distributed training (multi-GPU/node, GPU clusters), understanding associated memory and communication challenges. Real model-engineering experience in physics/simulation, vision, or LLM systems, with a focus on data-driven system improvement. Solid Python platform and backend foundations: API design, workflow orchestration, and practical Docker/Kubernetes/IaC. #J-18808-Ljbffr
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Highest Paying Tech Companies for Developers