about
AttentionRAG: Attention-Guided Context Pruning in Retrieval-Augmented Generation (arxiv.org)
2 points by PaulHoule on Mar 31, 2025 | hide | past | pdf | discuss on HN

In plain words: It turns the question into a guess-the-next-word task; its focus lands on one token, whose attention decides what to keep. It compressed retrieved text up to 6.3 times while answering about 10% better than the usual pruning tool that ignores the question.

Abstract

While RAG demonstrates remarkable capabilities in LLM applications, its effectiveness is hindered by the ever-increasing length of retrieved contexts, which introduces information redundancy and substantial computational overhead. Existing context pruning methods, such as LLMLingua, lack contextual awareness and offer limited flexibility in controlling compression rates, often resulting in either insufficient pruning or excessive information loss. In this paper, we propose AttentionRAG, an attention-guided context pruning method for RAG systems. The core idea of AttentionRAG lies in its attention focus mechanism, which reformulates RAG queries into a next-token prediction paradigm. This mechanism isolates the query's semantic focus to a single token, enabling precise and efficient attention calculation between queries and retrieved contexts. Extensive experiments on LongBench and Babilong benchmarks show that AttentionRAG achieves up to 6.3$\times$ context compression while outperforming LLMLingua methods by around 10\% in key metrics.

Yixiong Fang, Tianran Sun, Yuling Shi, Xiaodong Gu
arXiv:2503.10720 · cs.CL, cs.AI · submitted Mar 13, 2025 · updated Oct 27, 2025
abstract · pdf · html

add comment on HN