about
DeepSeek Native Sparse Attention (arxiv.org)
16 points by bandwitch on Feb 18, 2025 | hide | past | pdf | 1 comment on HN

In plain words: Instead of every token checking every other token, it reads a compressed summary of the text plus a chosen set of the most relevant tokens, and is trained from scratch this way. It matched or beat full attention and ran much faster on 64,000-token sequences.

Abstract · Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Long-context modeling is crucial for next-generation language models, yet the high computational cost of standard attention mechanisms poses significant computational challenges. Sparse attention offers a promising direction for improving efficiency while maintaining model capabilities. We present NSA, a Natively trainable Sparse Attention mechanism that integrates algorithmic innovations with hardware-aligned optimizations to achieve efficient long-context modeling. NSA employs a dynamic hierarchical sparse strategy, combining coarse-grained token compression with fine-grained token selection to preserve both global context awareness and local precision. Our approach advances sparse attention design with two key innovations: (1) We achieve substantial speedups through arithmetic intensity-balanced algorithm design, with implementation optimizations for modern hardware. (2) We enable end-to-end training, reducing pretraining computation without sacrificing model performance. As shown in Figure 1, experiments show the model pretrained with NSA maintains or exceeds Full Attention models across general benchmarks, long-context tasks, and instruction-based reasoning. Meanwhile, NSA achieves substantial speedups over Full Attention on 64k-length sequences across decoding, forward propagation, and backward propagation, validating its efficiency throughout the model lifecycle.

Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, Y. X. Wei, Lean Wang, Zhiping Xiao, Yuqing Wang, Chong Ruan, et al.
arXiv:2502.11089 · cs.CL, cs.AI, cs.LG · submitted Feb 16, 2025 · updated Feb 27, 2025
abstract · pdf · html

add comment on HN
Also discussed: Mar 2025 (2 points, 0 comments) · Feb 2025 (2 points, 0 comments) · Feb 2025 (4 points, 2 comments) · Feb 2025 (15 points, 2 comments)

Sparse attention essentially combines 3 types of attention optimizations:

1. Compression of the query input vectors to reduce the size of the KV cache

2. Selectively computing uncompressed attention on a subset of tokens based on the compressed blocks with the highest attention scores

3. Using sliding window for local attention at full resolution

> Both Full Attention and sparse attention models are pretrained on 270⁢B tokens of 8⁢k-length texts, followed by continued training and supervised fine-tuning on 32⁢k-length texts with YaRN to achieve long-context adaptation. Both models are trained to full convergence to ensure fair comparison.

> our experiments adopt a backbone combining Grouped-Query Attention (GQA) and Mixture-of-Experts (MoE), featuring 27⁢B total parameters with 3⁢B active parameters

Evaluated on MMLU, MMLU-PRO, CMMLU, BBH, GSM8K, MATH, DROP, MBPP, and HumanEval. NSA outperforms full attention on 7/9.

Beats out H2O, InfLLM, Quest, Exact-Top, and full attention on LongBench

Perfect retrieval on 64k needle-in-a-haystack

The CoT eval is less convincing, but outperforms the FA on AIME24.

Training speed of 2-9x vs. FlashAttention

Decoding speedup of 4-12x vs. full attention ["expected"? Didn't see comparison to other attention mechanisms]