about
Multimodal Neural Databases (arxiv.org)
2 points by srijansriv on Mar 19, 2024 | hide | past | pdf | 1 comment on HN

In plain words: A system that answers database-style questions—like counting or filtering—across mixed text and image collections at scale, where ordinary search tools only find similar items. Compared with today's multimodal models, it handles these multi-step queries far better.

Abstract

The rise in loosely-structured data available through text, images, and other modalities has called for new ways of querying them. Multimedia Information Retrieval has filled this gap and has witnessed exciting progress in recent years. Tasks such as search and retrieval of extensive multimedia archives have undergone massive performance improvements, driven to a large extent by recent developments in multimodal deep learning. However, methods in this field remain limited in the kinds of queries they support and, in particular, their inability to answer database-like queries. For this reason, inspired by recent work on neural databases, we propose a new framework, which we name Multimodal Neural Databases (MMNDBs). MMNDBs can answer complex database-like queries that involve reasoning over different input modalities, such as text and images, at scale. In this paper, we present the first architecture able to fulfill this set of requirements and test it with several baselines, showing the limitations of currently available models. The results show the potential of these new techniques to process unstructured data coming from different modalities, paving the way for future research in the area. Code to replicate the experiments will be released at https://github.com/GiovanniTRA/MultimodalNeuralDatabases

Giovanni Trappolini, Andrea Santilli, Emanuele Rodolà, Alon Halevy, Fabrizio Silvestri
arXiv:2305.01447 · cs.MM, cs.CL, cs.CV, cs.DB, cs.IR · submitted May 2, 2023
abstract · pdf · html

add comment on HN

This, specially the nlp part, strikes me as a primitive version of the RAG-LLM pair