about
Atlas: Learning to Optimally Memorize the Context at Test Time (arxiv.org)
43 points by famouswaffles on May 31, 2025 | hide | past | pdf | 4 comments on HN

In plain words: A new memory module stores context by tuning its memory with current and past tokens, not just the newest one like models that keep a fixed-size running summary. It beat Transformers and those models, with 80% higher accuracy at 10 million tokens of context.

Abstract · ATLAS: Learning to Optimally Memorize the Context at Test Time

Transformers have been established as the most popular backbones in sequence modeling, mainly due to their effectiveness in in-context retrieval tasks and the ability to learn at scale. Their quadratic memory and time complexity, however, bound their applicability in longer sequences and so has motivated researchers to explore effective alternative architectures such as modern recurrent neural networks (a.k.a long-term recurrent memory module). Despite their recent success in diverse downstream tasks, they struggle in tasks that requires long context understanding and extrapolation to longer sequences. We observe that these shortcomings come from three disjoint aspects in their design: (1) limited memory capacity that is bounded by the architecture of memory and feature mapping of the input; (2) online nature of update, i.e., optimizing the memory only with respect to the last input; and (3) less expressive management of their fixed-size memory. To enhance all these three aspects, we present ATLAS, a long-term memory module with high capacity that learns to memorize the context by optimizing the memory based on the current and past tokens, overcoming the online nature of long-term memory models. Building on this insight, we present a new family of Transformer-like architectures, called DeepTransformers, that are strict generalizations of the original Transformer architecture. Our experimental results on language modeling, common-sense reasoning, recall-intensive, and long-context understanding tasks show that ATLAS surpasses the performance of Transformers and recent linear recurrent models. ATLAS further improves the long context performance of Titans, achieving +80\% accuracy in 10M context length of BABILong benchmark.

Ali Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri, Yuan Deng, Peilin Zhong, Meisam Razaviyayn, Vahab Mirrokni
arXiv:2505.23735 · cs.CL, cs.AI · submitted May 29, 2025
abstract · pdf · html

add comment on HN

It seems like there’s been a lot of progress here, but it also seems like there’s an elephant in the room that RNNs will _always_ have worse memory than self-attention since the latter always has complete access to the full context. We pay for that in other ways, but it seems like the unstated hypothesis of RNNs is that we believe in the long run that RNNs will be “good enough” and their other performance benefits will eventually prevail. I’m not convinced that humanity will ever sink comparable resources into optimizing this family of models that has gone into Transformers to make them practical at the scale they run today.
The goal here is not to replace transformers but combine them with RNN so you get both good short-term memory (self-attention) and much improved long-term memory (ATLAS recurrent memory).

"Empirically, our models—OmegaNet, Atlas, DeepTransformers, and Dot—achieve consistent improvements over Transformers and recent RNN variants across diverse benchmarks."

The Titans papers and the Test-Time Training papers (https://arxiv.org/abs/2407.04620) both have the same premise - models should "learn" from their context rather than memorize them. Very promising direction!
Damn thought this was about like school tests, which is something I did a lot, try to memorize everything just before the test. Guess I got baited into clicking LLM article again