about
Twilight: Adaptive Attention Sparsity with Hierarchical Top-$P$ Pruning (arxiv.org)
2 points by PaulHoule on Feb 26, 2025 | hide | past | pdf | discuss on HN

In plain words: It keeps only the words a model pays attention to, stopping once their weight hits a set share, so the budget adapts. Added to sparse-attention tricks, it pruned up to 98% of redundant tokens and sped up long-context decoding 3.9 times without losing accuracy.

Abstract · Twilight: Adaptive Attention Sparsity with Hierarchical Top-$p$ Pruning

Leveraging attention sparsity to accelerate long-context large language models (LLMs) has been a hot research topic. However, current algorithms such as sparse attention or key-value (KV) cache compression tend to use a fixed budget, which presents a significant challenge during deployment because it fails to account for the dynamic nature of real-world scenarios, where the optimal balance between accuracy and efficiency can vary greatly. In this paper, we find that borrowing top-$p$ sampling (nucleus sampling) to sparse attention can surprisingly achieve adaptive budgeting. Based on this, we propose Twilight, a framework to bring adaptive sparsity to any existing sparse attention algorithm without sacrificing their accuracy. Empirical results show that Twilight can adaptively prune at most 98% of redundant tokens, leading to $15.4\times$ acceleration in self-attention operations and $3.9\times$ acceleration in end-to-end per token latency in long context LLM decoding.

Chaofan Lin, Jiaming Tang, Shuo Yang, Hanshuo Wang, Tian Tang, Boyu Tian, Ion Stoica, Song Han, Mingyu Gao
arXiv:2502.02770 · cs.LG, cs.CL · submitted Feb 4, 2025 · updated Nov 4, 2025
abstract · pdf · html · To appear on NeurIPS 2025 (spotlight)

add comment on HN