about
HMT: Hierarchical Memory Transformer for Long Context Language Processing (arxiv.org)
87 points by jasondavies on May 17, 2024 | hide | past | pdf | 6 comments on HN

In plain words: It layers memory like a brain, keeping text in tiers and pulling back relevant details so small models handle long inputs. It matched or beat long-context models on language modeling, question answering, and summarizing, with up to 57 times fewer parameters and less memory.

Abstract · HMT: Hierarchical Memory Transformer for Efficient Long Context Language Processing

Transformer-based large language models (LLM) have been widely used in language processing applications. However, due to the memory constraints of the devices, most of them restrict the context window. Even though recurrent models in previous works can memorize past tokens to enable unlimited context and maintain effectiveness, they have ``flat'' memory architectures. Such architectures have limitations in selecting and filtering information. Since humans are good at learning and self-adjustment, we believe that imitating brain memory hierarchy is beneficial for model memorization. Thus, we propose the Hierarchical Memory Transformer (HMT), a novel framework that facilitates a model's long-context processing ability by imitating human memorization behavior. Leveraging memory-augmented segment-level recurrence, we organize the memory hierarchy by preserving tokens from early input segments, passing memory embeddings along the sequence, and recalling relevant information from history. Evaluating general language modeling, question-answering tasks, and the summarization task, we show that HMT consistently improves the long-context processing ability of existing models. Furthermore, HMT achieves a comparable or superior generation quality to long-context LLMs with $2 \sim 57\times$ fewer parameters and $2.5 \sim 116\times$ less inference memory, significantly outperforming previous memory-augmented models. Code on Github: https://github.com/OswaldHe/HMT-pytorch.

Zifan He, Yingqi Cao, Zongyue Qin, Neha Prakriya, Yizhou Sun, Jason Cong
arXiv:2405.06067 · cs.CL, cs.LG · submitted May 9, 2024 · updated Feb 6, 2025
abstract · pdf · html · NAACL 2025 Main Conference

add comment on HN

Code: https://github.com/OswaldHe/HMT-pytorch

This looks really interesting. I've added the paper to my reading list and look forward to playing with the code. I'm curious to see what kinds of improvements we can get by agumenting Transformers and other generative sequence models with this and other mechanisms implementing hierarchical memory.[a]

Shouldn't the authors cite the work by Jeff Hawkins et al at Numenta? Hawkins has been proposing AI models with hierarchical temporal memory for a long time.[b] I can't help but wonder if there is a way, somehow, to incorporate his work and ideas in Transformers and other generative sequence models.

We sure live in interesting times!

---

[a] In the past, I've experimented with mechanisms that add memory to Transformers, but never with hierarchy.

[b] https://en.wikipedia.org/wiki/Hierarchical_temporal_memory

I thought Hawkins's book "On Intelligence" was amazing. It's a bit wild how close things have followed to the direction he laid out.
That's kinda hilarious, because I think the book was exactly wrong in its predictions. This can be evidenced by the continuous failures of his AI company, Numenta.
Well both things can be true at the same time.

I agree Numenta missed the boat, but it doesn't mean that book wasn't prescient. Numenta just didn't get there first, wasn't a fast follower, and blew a huge lead in AI. It may still end up the HTM is one of the final state solution,s but they are so far from ever being able to capitalize on it that it is unlikely to even matter if they invented the concept.

Does it relate to applications in time series? How does the hierarchy play a role in Transformers?
Not really. But questions about the work should be asked to the authors. I've only skimmed the paper.