Research Scientist - Computer Vision
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+3 more
Job description
We are seeking a Research Scientist specialising in computer vision and multimodal AI to join a leading research and development team in London.
This permanent, full-time role focuses on advancing next-generation artificial intelligence across areas such as multimodal understanding and generation, vision-language models, large-scale representation learning, and embodied intelligence.
You will work alongside an international team of researchers and engineers, with access to substantial computing resources and modern AI infrastructure. The role is based in King’s Cross, London.
Key Responsibilities
Advanced Model Research
- Design and develop Vision Transformer and multimodal large-model architectures with improved reasoning, efficiency, and scalability.
- Advance multimodal alignment, representation learning, and long-context modelling.
- Research scalable training techniques for large multimodal models.
- Improve model architectures to strengthen generalisation, robustness, and overall performance.
- Explore new approaches across multimodal understanding, generation, and reasoning.
Multimodal Data Development
- Process large-scale multimodal datasets spanning images, video, audio, and text.
- Build data pipelines for cleaning, filtering, annotation, validation, and quality control.
- Construct and maintain reproducible datasets with clear versioning and documentation.
- Optimise data mixtures, sampling methods, and curriculum strategies for model training.
- Improve dataset quality through evaluation results and feedback-driven curation.
- Develop methods for identifying low-quality, duplicated, biased, or uninformative data.
Large Multimodal Model Systems
- Build and improve distributed training systems for large-scale multimodal models.
- Optimise GPU utilisation, cluster efficiency, resource allocation, and workload scheduling.
- Develop scalable training frameworks and reusable research infrastructure.
- Engineer training, inference, evaluation, and serving systems.
- Improve the scalability, reliability, stability, and performance of model development pipelines.
- Work closely with infrastructure and platform teams to resolve system-level bottlenecks.
Research Application and Delivery
- Apply multimodal capabilities to intelligent assistants, content generation, and related AI applications.
- Translate research outcomes into production-ready systems and user-facing features.
- Collaborate with product, engineering, and research teams to deploy, evaluate, and iterate models.
- Communicate research findings through technical reports, presentations, publications, and demonstrations.
Requirements
- Bachelor’s degree or higher in Computer Science, Mathematics, Statistics, Engineering, or another relevant technical discipline.
- Strong Python programming skills.
- Hands-on experience with PyTorch or comparable deep learning frameworks.
- Strong ability to develop, implement, and evaluate machine learning algorithms.
- Solid mathematical reasoning and problem-solving skills.
- Good understanding of modern computer vision or multimodal machine learning methods.
- Ability to collaborate effectively across research, engineering, product, and infrastructure teams.
- Strong written and verbal communication skills.
- Self-motivated, resilient, and comfortable working on complex research problems with a high degree of technical uncertainty., * Master’s degree or PhD in Computer Science, Artificial Intelligence, Machine Learning, Computer Vision, or a related field.
- Publications at leading artificial intelligence or computer vision conferences, including CVPR, ICCV, ECCV, NeurIPS, ICML, or ICLR.
- Experience pre-training, fine-tuning, or evaluating large-scale vision or multimodal models.
- Experience with Vision Transformers, vision-language models, multimodal large language models, or generative vision systems.
- Familiarity with distributed training, mixed-precision training, model parallelism, or large-scale GPU clusters.
- Contributions to high-impact open-source projects in computer vision, natural language processing, multimodal AI, or machine learning systems.
- Research or internship experience within a recognised technology company, research laboratory, or academic institution.
- Experience translating research prototypes into scalable production systems.
Additional Skills
- Strong understanding of representation learning, attention mechanisms, transformers, and generative modelling.
- Experience working with large multimodal datasets and data-quality pipelines.
- Familiarity with model evaluation, benchmarking, ablation studies, and experiment reproducibility.
- Ability to identify research opportunities and independently drive projects from initial concept to validated outcome.
- Interest in advancing the capabilities, efficiency, and reliability of next-generation multimodal AI systems.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on eu-recruit.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What Are Large Language Models?
What Industries Outside of AI Are Hiring The Most AI Experts?
Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production