Multimodal Data Engine Expert

Huawei R&D Sites in Belgium and the Netherlands
Amsterdam, Netherlands
1 day ago
Apply on www.adzuna.nl
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

C (Programming Language) Java (Programming Language) Artificial Intelligence Data Analysis Computer Vision Big Data C++ (Programming Language) Computer Programming Databases Data Transmissions Information Engineering Data Infrastructure
+24 more
Data Systems DevOps Microprocessors Distributed Data Store Distributed Systems Python (Programming Language) Meta-Data Management Natural Language Processing Open Source Technology Distributed Caching Unstructured Data Reinforcement Learning Data Processing Graphics Processing Unit (GPU) Data Storage Technologies Cloud Platform System High Performance Computing Large Language Models Indexer Build Management Core Data Data Management Virtual Agents Serverless Computing

Job description

We are seeking a Multimodal Data Engine Expert to lead the development of next-generation data infrastructure for the AI Agent era.

The successful candidate will drive the architectural evolution and core technology development of a next-generation multimodal intelligent data platform, with a focus on heterogeneous compute resource pooling, multimodal computing engines, vector storage systems, high-performance distributed caching, AI4DB, and Data Agent technologies.

This role will be responsible for building end-to-end capabilities for massive-scale heterogeneous and unstructured data, spanning resource scheduling, computation, retrieval, storage, and intelligent autonomous management. The goal is to establish core data infrastructure optimized for large language models and AI Agent ecosystems, while driving technological innovation and practical implementation at the intersection of Data and AI., Design and build a unified scheduling and management platform for heterogeneous computing resources, including CPUs, GPUs, and NPUs. Develop serverless resource pooling and management capabilities across multiple computing engines, continuously optimize heterogeneous resource scheduling performance, and improve overall resource utilization.

  1. Multimodal Computing Engine Development Develop and optimize computing engines for large-scale unstructured and multimodal data, enabling efficient processing and analysis of text, images, video, and other data types. Design next-generation multimodal query optimizers and hybrid execution engines, and support high-performance retrieval and processing of vector, textual, geospatial, and other data.

  2. Multimodal Vector Indexing and Management Systems Design and develop unified systems for multimodal vector retrieval, storage, metadata management, and access control. Integrate with open-source ecosystems and enable high-performance, intelligent storage optimization, data management, and index lifecycle management.

  3. Distributed Caching Services for Multimodal Data Build high-performance distributed caching services for multimodal data platforms. Provide high-speed near-compute caching for computing engines and efficient east-west data transfer across engines, improving the overall performance and efficiency of multimodal data processing pipelines.

  4. AI-Powered Agent Development Leverage emerging AI4DB (AI for Databases) and LLM Agent technologies to build next-generation intelligent data platforms with autonomous decision-making capabilities. Develop intelligent Agents and Skills that enable end-to-end automation and intelligence across workload development, data storage, data analysis, operations, and system maintenance.

Requirements

  1. Strong proficiency in programming languages such as C, C++, Python, and Java.
  2. Strong R&D background and technical expertise in areas such as database systems, big data systems, distributed systems, and high-performance computing. Research or hands-on development experience with core system components, such as query optimizers, execution engines, storage engines, or distributed storage systems, is highly preferred.
  3. Strong understanding of the architecture, characteristics, and performance considerations of heterogeneous computing resources, including CPUs, GPUs, and NPUs. Hands-on experience in heterogeneous resource management and scheduling, with the ability to enable unified resource management and efficient resource utilization.
  4. Experience in cloud computing platform development and maintenance, as well as DevOps-related engineering practices.
  5. Experience with LLM fine-tuning and reinforcement learning, multimodal data processing technologies such as NLP, computer vision, and time-series analysis, and the integration of AI technologies with database or data systems is highly preferred.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.nl
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:36 min

Analyzing limitations with PostgreSQL bitmap heap scans

Dharin Shah Dharin Shah · World Congress 2025

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

3:50 min

Combining distinct processing modalities to elevate network intelligence

Ekaterina Sirazitdinova · LIVE

Videos

See all

Related articles

See all