In plain words: Instead of letting each token look only at nearby neighbors, this design adds a fixed long-range lookup as new text arrives, so distant links stay reachable. It beat local-window attention on language and long-sequence tasks, nearly matching full attention quality at linear cost.
Abstract · $π$-Attention: Online Efficient Sparse Transformers for Long-Context Modeling
Sparse attention is crucial in long-context Transformers, which restricts each token to a limited neighborhood and thereby reduces the quadratic cost of full self-attention. Local windows capture nearby context effectively, yet they induce a receptive-field bottleneck for dependencies beyond the window, limiting long-range modeling under moderate depth. In this paper, we propose $π$-Attention, an \emph{online efficient} sparse attention operator: as tokens arrive, each step maintains a streaming working set of local neighbors plus a $π$-indexed long-range fetch, fused by an adaptive prior under a shared softmax. Rather than materializing a global sparse mask in advance, $π$-Attention computes attention on the live working set with hierarchy-aware IO. We analyze causal reachability and minimum depth under this online rule, and show per-step cost remains $\mathcal{O}(k)$. Experiments on language modeling, Long Range Arena, and efficiency profiling---across 4K--32K context lengths---show consistent gains over local-window and other sparse baselines, approaching dense attention quality at linear cost.
Pike D. Liu, Chang Liu, Yanxuan Yu
arXiv:2511.10696 · cs.CL, cs.AI · submitted Nov 12, 2025 · updated Aug 4, 2026
abstract · pdf · html