In plain words: Instead of a separate search engine plus a language model, one model learns the document collection, then writes out the relevant passages itself and judges which truly match the question. It beat standard search systems clearly and improved answers that rely on retrieved documents.
Abstract · Self-Retrieval: End-to-End Information Retrieval with One Large Language Model
The rise of large language models (LLMs) has significantly transformed both the construction and application of information retrieval (IR) systems. However, current interactions between IR systems and LLMs remain limited, with LLMs merely serving as part of components within IR systems, and IR systems being constructed independently of LLMs. This separated architecture restricts knowledge sharing and deep collaboration between them. In this paper, we introduce Self-Retrieval, a novel end-to-end LLM-driven information retrieval architecture. Self-Retrieval unifies all essential IR functions within a single LLM, leveraging the inherent capabilities of LLMs throughout the IR process. Specifically, Self-Retrieval internalizes the retrieval corpus through self-supervised learning, transforms the retrieval process into sequential passage generation, and performs relevance assessment for reranking. Experimental results demonstrate that Self-Retrieval not only outperforms existing retrieval approaches by a significant margin, but also substantially enhances the performance of LLM-driven downstream applications like retrieval-augmented generation.
Qiaoyu Tang, Jiawei Chen, Zhuoqun Li, Bowen Yu, Yaojie Lu, Cheng Fu, Haiyang Yu, Hongyu Lin, Fei Huang, Ben He, Xianpei Han, Le Sun, et al.
arXiv:2403.00801 · cs.IR, cs.AI, cs.CL · submitted Feb 23, 2024 · updated Nov 4, 2024
abstract · pdf · html · NeurIPS 2024 Camera-ready Version. Code: https://github.com/icip-cas/SelfRetrieval
> To accurately generate the exact passages in the given corpus, we employ a trie-based constrained decoding algorithm (Chen et al., 2020; Cao et al., 2021; Lu et al., 2021) in which the generated tokens can be constrained in the dynamic vocabulary. Specifically, instead of generating a token from the entire target vocabulary at each step, we use a prefix tree (trie) to constraint the target vocabulary and ensure that the generated content is within the corpus. During the construction of trie, we remove stop words from the initial token to improve semantic representation of the trie.