In plain words: The model periodically "sleeps": it runs extra offline passes over recent context to fold it into lasting memory, then clears its stored history so answering stays fast. More sleep passes improved accuracy most on the hardest reasoning problems, where standard transformers failed.
Abstract · Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference
Transformer-based large language models are increasingly used for long-horizon tasks; however, their attention mechanism scales poorly with context length. To handle this, we study a sleep-like consolidation mechanism in which a model periodically converts recent context into persistent fast weights before clearing its key-value cache. During sleep, the model performs $N$ offline recurrent passes over the accumulated context and updates the fast weights in its state-space model (SSM) blocks through a learned local rule. During inference, this shifts extra computation to sleep while preserving the latency of wake-time prediction. We test our method on controlled synthetic tasks, including cellular automata and multi-hop graph retrieval, as well as a realistic math reasoning task, on which a regular transformer as well as SSM-attention hybrid models fail. We then show that increasing sleep duration $N$ for our models improves performance, with the largest gains on examples that require deeper reasoning.
Sangyun Lee, Sean McLeish, Tom Goldstein, Giulia Fanti
arXiv:2605.26099 · cs.CL, cs.AI · submitted May 25, 2026 · updated Jun 5, 2026
abstract · pdf · html
Essentially it goes "You know how your model can remember its training data? Well, what if you treated its recent context like more training data and updated (some of) the weights using (mostly) the same process used to train it?"
The end result is very good at remembering things but also really good at adapting to new unseen distributions.
[1]https://arxiv.org/abs/2512.23675