about
LM2: Large Memory Models (arxiv.org)
110 points by fzliu on Feb 13, 2025 | hide | past | pdf | 30 comments on HN

In plain words: A text model keeps a separate memory store that reads and updates alongside the normal flow, to hold facts across long passages. On long-context question answering it beat a rival memory system by 37.1%, and stayed 5% better than the plain model on general tests.

Abstract

This paper introduces the Large Memory Model (LM2), a decoder-only Transformer architecture enhanced with an auxiliary memory module that aims to address the limitations of standard Transformers in multi-step reasoning, relational argumentation, and synthesizing information distributed over long contexts. The proposed LM2 incorporates a memory module that acts as a contextual representation repository, interacting with input tokens via cross attention and updating through gating mechanisms. To preserve the Transformers general-purpose capabilities, LM2 maintains the original information flow while integrating a complementary memory pathway. Experimental results on the BABILong benchmark demonstrate that the LM2model outperforms both the memory-augmented RMT model by 37.1% and the baseline Llama-3.2 model by 86.3% on average across tasks. LM2 exhibits exceptional capabilities in multi-hop inference, numerical reasoning, and large-context question-answering. On the MMLU dataset, it achieves a 5.0% improvement over a pre-trained vanilla model, demonstrating that its memory module does not degrade performance on general tasks. Further, in our analysis, we explore the memory interpretability, effectiveness of memory modules, and test-time behavior. Our findings emphasize the importance of explicit memory in enhancing Transformer architectures.

Jikun Kang, Wenqi Wu, Filippos Christianos, Alex J. Chan, Fraser Greenlee, George Thomas, Marvin Purtorab, Andy Toulis
arXiv:2502.06049 · cs.CL, cs.AI · submitted Feb 9, 2025
abstract · pdf · html

add comment on HN

(Probably bc I'm dumb) I'm very confused by this paper. The dimensions are all over the place: first they say M is a N x d x d matrix, then it becomes N x d. And then they are trying to scale M with g_out and add it to E_attn which is a T x d matrix??? Are the gates scalars or vectors or matrices? If they are matrices then the dimensions also don't line up to M
You're not dumb. I think it's just poorly written and full of errors.
Missed opportunity to call them LMM and enjoy the onslaught of typos
Some Chinese researchers have worked on copper nanotubes and abbreviated them as CuNTs [0].

[0]: https://www.researchgate.net/publication/260800015_Structura...

Or Models with Large Memory. MLMs :)
This seems to be the most appropriate
I can't wait to enhance my STT-RAG-LLM-TTS with some LMMs
You forgot SOTA, that one is flying around a lot.
Also left out SoDoSoPa
Wouldn't it be STT-LLM-RAG-TTS
They're still language models

LMLM

Just as Graph RAG should have just been GAG.
That's why they named it LM2 ..
Seems like LM² would have made more sense in that case. Or L²M² Or LLMM
The question is whether L and M commute.
Because that makes it harder to type out
(LM)²
The largest model they tested is 1.7B.
Do they believe that it's likely to scale as high as "standard" LLMs? Or is this another state space machine moment where it turns out to be equivalent to them?
The concept of a memory module, separate from a prediction engine, makes sense to me, our brains might operate like that (short vs long term memory).

Many research groups have been trying to implement this idea, recent paper from Meta comes to mind: https://arxiv.org/html/2412.09764

If they believe it’s likely to scale, they will try to scale it, so we will either hear about it in a few months, or we won’t.

Table 2's results are interesting. If the paper is to be believed, just adding the memory model seems to improve reasoning tasks across the board.

That said, I do wonder if this a bit of mirage. At 1.7B parameters, they are 3 orders of magnitude down from 4o (well that isn't completely fair, I don't know what the average 'expert' size is in 4o, but I doubt the authors are doing mixture of experts at only 1.7B). A model can 'memorize' way more shit with that many parameters.

This immediately makes my mind bring up Hopfield networks https://arxiv.org/abs/2008.02217

when I worked with them circa 2012 they were practically toys. Maybe we are in a better place now?

RNN with extra steps?
There are many papers that use a recurrence across sub-sequences and attention within sub-sequences. Google did this with Infini-Attention and one of the variants from the Titans paper. However, I think the earliest example of this is Transformer-XL.
Isn't that all of modern AI?
Transformers are completely unlike RNNs.
There are some interesting connections between them. If you remove the softmax from the attention formula, you end up with linear attention, which has a recurrent form.

I haven't read it, but the Mamba 2 paper claims to establish a stronger connection.

* If you remove the softmax from the attention formula, you end up with linear attention*

Sorry, what?

Here is a paper explaining it: https://arxiv.org/abs/2006.16236
GitHub link in the paper is a 404 - private repo?