about
Infinite Retrieval: Attention enhanced LLMs in long-context processing (arxiv.org)
37 points by TaurenHunter on Mar 1, 2025 | hide | past | pdf | 7 comments on HN

In plain words: A small model's own attention scores are used to hunt down the relevant pieces of a huge input, so it can answer questions from text of any length without extra training or a separate search tool. It found a hidden fact in 1 million tokens with 100% accuracy, beating larger models and other approaches.

Abstract · Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing

Limited by the context window size of Large Language Models(LLMs), handling various tasks with input tokens exceeding the upper limit has been challenging, whether it is a simple direct retrieval task or a complex multi-hop reasoning task. Although various methods have been proposed to enhance the long-context processing capabilities of LLMs, they either incur substantial post-training costs, or require additional tool modules(e.g.,RAG), or have not shown significant improvement in realistic tasks. Our work observes the correlation between the attention distribution and generated answers across each layer, and establishes the attention allocation aligns with retrieval-augmented capabilities through experiments. Drawing on the above insights, we propose a novel method InfiniRetri that leverages the LLMs's own attention information to enable accurate retrieval across inputs of infinitely length. Our evaluations indicate that InfiniRetri achieves 100% accuracy in the Needle-In-a-Haystack(NIH) test over 1M tokens using a 0.5B parameter model, surpassing other method or larger models and setting a new state-of-the-art(SOTA). Moreover, our method achieves significant performance improvements on real-world benchmarks, with a maximum 288% improvement. In addition, InfiniRetri can be applied to any Transformer-based LLMs without additional training and substantially reduces inference latency and compute overhead in long texts. In summary, our comprehensive studies show InfiniRetri's potential for practical applications and creates a paradigm for retrievaling information using LLMs own capabilities under infinite-length tokens. Code will be released in link.

Xiaoju Ye, Zhichun Wang, Jingyuan Wang
arXiv:2502.12962 · cs.CL · submitted Feb 18, 2025
abstract · pdf · html · 21 pages

add comment on HN

This paper highlights something that should have been obvious: prediction and retrieval are two sides of the same coin. To predict effectively, you must first identify what's relevant. What's remarkable is that a 0.5B parameter model can perform perfect retrieval over 1M tokens when its natural attention patterns are leveraged properly.

It raises an interesting question: what if we designed architectures explicitly around retrieval capabilities? Transformer architectures were designed for prediction, and retrieval emerged as a byproduct. What would an architecture optimized specfically for retrieval look like?

A lot of money has been spent on building out large-scale RAG systems. If the performance improvements promised by the paper are real, the ramifications will be huge. Exciting to see that the authors are promising to release their code - it will be fun to how this model performs on consumer hardware.

I think this could be expanded further. You can convert attention traces to knowledge graph with arbitrary and/or dynamic density. Traversing it can also be exotic – zooming in/expanding details at arbitrary points during traversal. With common format you can create topic/knowledge trace packs that could be shared, merged (subtracted?) etc.
I read through the paper, and I found the insights to be excellent.

However, regarding the practical implementation, the paper assumes that the questions will be available in advance. For each question, it requires calculating attention scores between the question and the context chunks, which makes it impractical as a replacement for Retrieval-Augmented Generation (RAG). For instance, if there are 1,000 documents, each with 10 chunks, it would be infeasible to compute attention scores between 10,000 chunks and a user query every time.

Am I correct in thinking that RAG, or SFT, would still be needed to introduce unseen context to the model.
Using attention for the retrieval of relevant information seems super intuitive. Only feed the model what it deems relevant. Curious about the scenarios where this mechanism misses relevant information.
Do I understand right this requires access to internals of the LLM and can not be used with todays models behind an API like ChatGPT or Claude?
Innovation that would be applicable to open weight models running locally only would be awesome.