about
State-space models can learn in-context by gradient descent (arxiv.org)
86 points by dsalaj on Oct 26, 2024 | hide | past | pdf | 58 comments on HN

In plain words: A linear recurrent network gets an extra step multiplying inputs in its memory, letting it run gradient descent on a simple predictor as it reads a sequence. Trained networks matched the predicted learning behavior on synthetic tasks and helped on long-range and language tasks.

Abstract · Learning in the Recurrent State: Gradient Descent with Linear Recurrent Networks

Linear recurrent networks (LRNNs) offer linear-time sequence modeling, but standard recurrent updates do not directly expose the supervised products needed for in-context gradient descent. We propose a sufficient constructive inductive bias for LRNNs: equip a diagonal recurrent state with multiplicative readout and a short sliding-window cross-product self-attention update. The resulting architecture, Gradient-based Recurrent In-context Learner (GRIL), can implement minibatch gradient descent on a task-specific linear predictor during a single forward pass. The same design extends to multi-step updates and cross-entropy classification, with a limited MLP-based extension to non-linear regression. Empirically, trained GRILs recover the behavior and parameters predicted by the construction on synthetic ICL tasks, and the same architectural bias yields useful performance on Long Range Arena and language modelling. These results present windowed cross-product self-attention as a practical, testable inductive bias for LRNNs that learn in context through gradient-descent-like updates.

Yudou Tian, Neeraj Mohan Sushma, Harshvardhan Mestha, Nicolo Colombo, David Kappel, Anand Subramoney
arXiv:2410.11687 · cs.LG, cs.AI, cs.NE · submitted Oct 15, 2024 · updated Jun 15, 2026
abstract · pdf · html · 28 pages, 11 figures

add comment on HN
Also discussed: Oct 2024 (2 points, 0 comments)

The discussion couldn't be loaded right now. Read it on Hacker News.