about
Kalman Delta Networks (arxiv.org)
3 points by E-Reverance 24 days ago | hide | past | pdf | discuss on HN

In plain words: A long-context model keeps a fixed-size memory of past words and tracks how sure it is about each stored association, so new information counts more when the memory is shaky. At two sizes, it beat top similar models on word-prediction error and task accuracy.

Abstract · Kalman Delta Networks: Uncertainty-aware Associative Memory

Linear attention enables efficient long-context inference by compressing token history into a fixed-size recurrent memory. This compression makes each update a trade-off between incorporating new information and preserving useful associations. Models such as DeltaNet, Gated DeltaNet, and KDA predict write strength from the current token representation, without explicitly tracking uncertainty in the stored memory. Yet this uncertainty matters: a new observation should have greater influence when the existing association is uncertain and less when it is already well supported. We introduce Kalman Delta Networks (KDNs), a family of linear-attention models that explicitly track memory uncertainty to guide each update. By formulating associative memory as a linear-Gaussian state-space model, KDNs propagate both the memory estimate and its uncertainty, using the Kalman gain to balance accumulated evidence against the reliability of new observations. This formulation also recovers standard delta-rule updates by replacing tracked covariance with a token-predicted isotropic surrogate. To support hardware-efficient training and inference, we derive Diagonal KDN and Isotropic KDN, which retain one uncertainty value per key channel and per head, respectively. Their uncertainty updates admit associative scans with logarithmic parallel depth, requiring only $O(d_k)$ and $O(1)$ auxiliary state per head. Across controlled pretraining at 750M and 1.3B parameters, both variants consistently improve perplexity and mean downstream accuracy over the evaluated state-of-the-art linear-attention baselines.

Ngoc Bui, Tinglin Huang, Rex Ying
arXiv:2609.07816 · cs.LG, cs.AI · submitted Sep 7, 2026 · updated Sep 29, 2026
abstract · pdf · html

add comment on HN