> Markdown version of [/jobs/ext/2425628-staff-research-engineer-scientific-computing-and-ml-physics-infrastructure](https://www.wearedevelopers.com/jobs/ext/2425628-staff-research-engineer-scientific-computing-and-ml-physics-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Research Engineer, Scientific Computing and ML/Physics Infrastructure - **Company:** Lila Sciences - **Location:** London, UK - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Systems Engineering, Profiling, Software Quality, Nvidia CUDA, Learning Management Systems, Linux, Distributed Systems, Fault Tolerance, General-Purpose Computing on Graphics Processing Units, Python (Programming Language), Machine Learning, Molecular Modeling, Scientific Computating, Software Engineering, Data Processing, Graphics Processing Unit (GPU), Pytorch, Large Language Models, Kubernetes, Slurm, Hardware Infrastructure, Data Pipelines, Docker - **Published:** August 3, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=8830f9bb1ed2220b ## About the Role * Strong software engineering skills in Python and experience working with ML, scientific computing, or simulation codebases. * Experience building, scaling, or operating distributed systems for research, ML, physics, simulation, or data-intensive workloads. * Practical knowledge of GPU computing, performance profiling, distributed execution, and failure modes in large-scale workloads. * Experience with PyTorch, JAX, CUDA-aware workflows, or related ML/scientific computing frameworks. * Practical knowledge of Linux, Docker or containers, dependency management, and reproducible development environments. * Experience with orchestration, scheduling, or distributed execution systems such as Kubernetes, Slurm, Ray, Flyte, Argo, or similar tools. * Ability to take prototype-quality research code and improve its architecture, scalability, reliability, and maintainability. * Strong debugging skills across code, environments, infrastructure, data pipelines, and compute clusters. * Ability to work directly with researchers, understand ambiguous technical needs, and convert them into robust engineering solutions. Bonus Points For * Familiarity with chemistry, computational biophysics, molecular simulation, computational chemistry, cheminformatics, or drug discovery workflows. * Experience with cloud GPU infrastructure, multi-cluster execution, or hybrid compute environments. * Experience building tools for LLM agents or automated research workflows. * Experience with workflow observability, checkpointing, retries, and fault-tolerant scientific workloads. * Experience with CI, testing, packaging, and release practices for research software. * Comfort supporting fast-moving research teams without over-engineering exploratory work. ## Description Lila Sciences is seeking a Research Engineer, Scientific Computing and ML/Physics Infrastructure to help turn promising research tools into robust, scalable systems. This role bridges research and production: you will work with scientists and ML researchers who can prototype useful tools, then help make those tools efficient, distributed, fault tolerant, and usable across Lila's compute environments. The Molecular Intelligence team is building ML and physics-based infrastructure for drug discovery, including biophysics workflows, computational chemistry tools, cofolding models, low-data learning systems, simulation workflows, and agent-usable scientific pipelines. We need an engineer who can improve code quality, architecture, GPU efficiency, cluster portability, and operational reliability without slowing down research velocity. What You'll Be Building * Take research tools, prototypes, and scientific workflows developed by scientists or academic-style researchers and make them scalable, efficient, and maintainable. * Collaborate directly with computational biophysics, computational chemistry, and machine learning scientists to turn research workflows into scalable agent-usable systems. * Build and support ML and physics infrastructure for model training, molecular simulation, data processing, and agent-executed scientific workflows. * Ensure workflows run reliably across multiple clusters and compute environments. * Improve GPU utilization, distributed execution, throughput, fault tolerance, and reproducibility for ML and scientific workloads. * Architect larger-scale systems around research code, including job orchestration, retry behavior, monitoring, artifact handling, and workflow traceability. * Optimize ML, physics, and pipeline code for performance and scalability. * Maintain development and execution environments across local, cloud, and GPU-based systems. * Package scientific tools into reusable services, workflows, or APIs that can be used by researchers, pipelines, and AI agents. * Partner with research, platform, and infrastructure teams to bridge exploratory scientific work with reliable engineering systems. * Document systems clearly and establish pragmatic engineering patterns for research teams. ## Related Videos - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)