> Markdown version of [/jobs/ext/1198355-multimodal-ml-engineer](https://www.wearedevelopers.com/jobs/ext/1198355-multimodal-ml-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Multimodal ML Engineer - **Company:** White Circle - **Location:** London, UK - **Experience:** Experienced - **Salary:** £120,000.0 - £250,000.0 - **Contract:** Permanent contract - **Skills:** Clean Code Principles, Audio Signal Processing, Noise Reduction, Distributed Computing Environment, Pytorch, Large Language Models, Deep Learning, Machine Learning Operations, Software Version Control, Data Pipelines, Data Generation - **Published:** July 7, 2026 - **Apply:** https://uk.indeed.com/viewjob?jk=8a2f19febe8ee2b6 ## About the Role * 3+ years training large-scale deep learning models in multimodal domains (vision-language, audio, speech, or acoustic) * Strong PyTorch skills with hands-on distributed training experience (DeepSpeed, FSDP, or similar) * Deep experience with multimodal architectures - you understand how vision/audio encoders, projectors, and LLMs fit together (LLaVA, Qwen-VL, InternVL, Audio Flamingo, Omni Qwen, Audio Qwen, Whisper, HuBERT, Conformer, or similar) * Hands-on with RLHF/alignment for multimodal: GRPO, DPO, reward modeling - not just for text * Experience with video and/or audio sequence modeling: temporal modeling, long-context processing, efficient attention, streaming inference * Track record of shipping models to production: you've hit latency targets and optimized inference, not just reported benchmark scores * Comfortable with large-scale multimodal dataset curation: image-text pairs, video-instruction data, audio preprocessing, augmentation, synthetic data generation * Familiar with MoE architectures and their tradeoffs for multimodal workloads * Strong engineering fundamentals: clean code, version control, testing, documentation A big plus: * Understanding of audio signal processing fundamentals (spectrograms, mel features, noise reduction) is a plus ## Description * Train and fine-tune large-scale multimodal models (vision-language, audio, speech) from scratch and from pretrained checkpoints * Extend models across modalities: image understanding, video temporal modeling, long-context processing, and streaming audio * Design and run experiments: architecture changes, data mixes, training recipes * Build and maintain multimodal data pipelines - from raw images, video, and audio recordings to training-ready datasets, including synthetic data generation * Train and optimize MoE architectures for efficient multimodal inference * Build alignment pipelines: SFT, DPO, GRPO, reward modeling - across modalities, not just text * Optimize models for production: quantization, distillation, batching, streaming and low-latency serving * Deploy models end-to-end: from research checkpoint to production serving * Define evaluation metrics and benchmarks that actually matter for the product: visual QA, spatial reasoning, video comprehension, speech and audio understanding ## Related Videos - [Getting Started with Machine Learning](https://www.wearedevelopers.com/videos/260-getting-started-with-machine-learning) - [Multimodal Generative AI Demystified](https://www.wearedevelopers.com/videos/829-multimodal-generative-ai-demystified) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Inside the Mind of an LLM](https://www.wearedevelopers.com/videos/1617-inside-the-mind-of-an-llm) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 138 - Are you secure about this?](https://www.wearedevelopers.com/magazine/486-dev-digest-138-are-you-secure-about-this)