Staff Research Engineer, Scientific Computing and ML/Physics Infrastructure

Lila Sciences
London, UK
about 1 month ago
Apply on uk.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Systems Engineering Profiling Software Quality Nvidia CUDA Learning Management Systems Linux Distributed Systems Fault Tolerance General-Purpose Computing on Graphics Processing Units Python (Programming Language)
+13 more
Machine Learning Molecular Modeling Scientific Computating Software Engineering Data Processing Graphics Processing Unit (GPU) Pytorch Large Language Models Kubernetes Slurm Hardware Infrastructure Data Pipelines Docker

Job description

Lila Sciences is seeking a Research Engineer, Scientific Computing and ML/Physics Infrastructure to help turn promising research tools into robust, scalable systems. This role bridges research and production: you will work with scientists and ML researchers who can prototype useful tools, then help make those tools efficient, distributed, fault tolerant, and usable across Lila’s compute environments.

The Molecular Intelligence team is building ML and physics-based infrastructure for drug discovery, including biophysics workflows, computational chemistry tools, cofolding models, low-data learning systems, simulation workflows, and agent-usable scientific pipelines. We need an engineer who can improve code quality, architecture, GPU efficiency, cluster portability, and operational reliability without slowing down research velocity.

What You’ll Be Building

  • Take research tools, prototypes, and scientific workflows developed by scientists or academic-style researchers and make them scalable, efficient, and maintainable.
  • Collaborate directly with computational biophysics, computational chemistry, and machine learning scientists to turn research workflows into scalable agent-usable systems.
  • Build and support ML and physics infrastructure for model training, molecular simulation, data processing, and agent-executed scientific workflows.
  • Ensure workflows run reliably across multiple clusters and compute environments.
  • Improve GPU utilization, distributed execution, throughput, fault tolerance, and reproducibility for ML and scientific workloads.
  • Architect larger-scale systems around research code, including job orchestration, retry behavior, monitoring, artifact handling, and workflow traceability.
  • Optimize ML, physics, and pipeline code for performance and scalability.
  • Maintain development and execution environments across local, cloud, and GPU-based systems.
  • Package scientific tools into reusable services, workflows, or APIs that can be used by researchers, pipelines, and AI agents.
  • Partner with research, platform, and infrastructure teams to bridge exploratory scientific work with reliable engineering systems.
  • Document systems clearly and establish pragmatic engineering patterns for research teams.

Requirements

  • Strong software engineering skills in Python and experience working with ML, scientific computing, or simulation codebases.
  • Experience building, scaling, or operating distributed systems for research, ML, physics, simulation, or data-intensive workloads.
  • Practical knowledge of GPU computing, performance profiling, distributed execution, and failure modes in large-scale workloads.
  • Experience with PyTorch, JAX, CUDA-aware workflows, or related ML/scientific computing frameworks.
  • Practical knowledge of Linux, Docker or containers, dependency management, and reproducible development environments.
  • Experience with orchestration, scheduling, or distributed execution systems such as Kubernetes, Slurm, Ray, Flyte, Argo, or similar tools.
  • Ability to take prototype-quality research code and improve its architecture, scalability, reliability, and maintainability.
  • Strong debugging skills across code, environments, infrastructure, data pipelines, and compute clusters.
  • Ability to work directly with researchers, understand ambiguous technical needs, and convert them into robust engineering solutions.

Bonus Points For

  • Familiarity with chemistry, computational biophysics, molecular simulation, computational chemistry, cheminformatics, or drug discovery workflows.
  • Experience with cloud GPU infrastructure, multi-cluster execution, or hybrid compute environments.
  • Experience building tools for LLM agents or automated research workflows.
  • Experience with workflow observability, checkpointing, retries, and fault-tolerant scientific workloads.
  • Experience with CI, testing, packaging, and release practices for research software.
  • Comfort supporting fast-moving research teams without over-engineering exploratory work.

About the company

Lila Sciences is building Scientific Superintelligence to solve humankind’s greatest challenges. We believe science is the most inspiring frontier for AI. Rather than hard-coding expert knowledge into tools, LILA builds systems that can learn for themselves.

LILA combines advanced AI models with proprietary AI Science Factory instruments into an operating system for science that executes the entire scientific method autonomously, accelerating discovery at unprecedented speed, scale, and impact across medicine, materials, and energy. Learn more at www.lila.ai.

Guided by our core values of truth, trust, curiosity, grit, and velocity, we move with startup speed while tackling problems of historic importance. If this sounds like an environment you’d love to work in, even if you don’t meet every qualification listed above, we encourage you to apply.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on uk.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all