about
Kimi Linear: An Expressive, Efficient Attention Architecture (arxiv.org)
6 points by birriel 334 days ago | hide | past | pdf | discuss on HN

In plain words: A new language model mixes standard attention with a cheaper memory-based layer that keeps a small running summary of past text, updated by fine-grained gates. Trained with the same recipe, it beat full attention on every task tested while storing 75% less of the past text in memory.

Abstract

We introduce Kimi Linear, a hybrid linear attention architecture that, for the first time, outperforms full attention under fair comparisons across various scenarios -- including short-context, long-context, and reinforcement learning (RL) scaling regimes. At its core lies Kimi Delta Attention (KDA), an expressive linear attention module that extends Gated DeltaNet with a finer-grained gating mechanism, enabling more effective use of limited finite-state RNN memory. Our bespoke chunkwise algorithm achieves high hardware efficiency through a specialized variant of the Diagonal-Plus-Low-Rank (DPLR) transition matrices, which substantially reduces computation compared to the general DPLR formulation while remaining more consistent with the classical delta rule. We pretrain a Kimi Linear model with 3B activated parameters and 48B total parameters, based on a layerwise hybrid of KDA and Multi-Head Latent Attention (MLA). Our experiments show that with an identical training recipe, Kimi Linear outperforms full MLA with a sizeable margin across all evaluated tasks, while reducing KV cache usage by up to 75% and achieving up to 6 times decoding throughput for a 1M context. These results demonstrate that Kimi Linear can be a drop-in replacement for full attention architectures with superior performance and efficiency, including tasks with longer input and output lengths. To support further research, we open-source the KDA kernel and vLLM implementations, and release the pre-trained and instruction-tuned model checkpoints.

Kimi Team, Yu Zhang, Zongyu Lin, Xingcheng Yao, Jiaxi Hu, Fanqing Meng, Chengyin Liu, Xin Men, Songlin Yang, Zhiyuan Li, Wentao Li, Enzhe Lu, et al.
arXiv:2510.26692 · cs.CL, cs.LG · submitted Oct 30, 2025 · updated Nov 1, 2025
abstract · pdf · html · Kimi Linear tech report

add comment on HN
Also discussed: Jul 2026 (306 points, 133 comments) · Jul 2026 (3 points, 0 comments)