> Markdown version of [/jobs/ext/2791052-senior-ml-performance-engineer-real-time-inference-scale](https://www.wearedevelopers.com/jobs/ext/2791052-senior-ml-performance-engineer-real-time-inference-scale). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior ML Performance Engineer - Real-Time Inference & Scale - **Company:** Odyssey - **Location:** Greater London, UK - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Machine Learning, Performance Tuning, Software Engineering, Pytorch - **Published:** September 8, 2026 - **Apply:** https://www.collegerecruiter.com/job/2840672798-senior-ml-performance-engineer-real-time-inference--scale ## About the Role Odyssey in Greater London seeks an experienced software engineer specializing in machine learning performance optimization. You will optimize models for real-time users, design distributed training strategies, and work with elite ML researchers. Candidates should have at least 8 years of software engineering experience, deep insights into machine learning architectures, and proficiency in PyTorch and NVIDIA optimization. This position offers autonomy in technical decisions and a chance to work with cutting-edge technology. #J-18808-Ljbffr ## Related Videos - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Practical performance tuning for Serverless Java on AWS](https://www.wearedevelopers.com/videos/2075-practical-performance-tuning-for-serverless-java-on-aws) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [30 Golden Rules of Deep Learning Performance](https://www.wearedevelopers.com/videos/11-30-golden-rules-of-deep-learning-performance) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Software Engineer Salary London](https://www.wearedevelopers.com/magazine/252-software-engineer-salary-london) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk)