In plain words: It layers memory like a brain, keeping text in tiers and pulling back relevant details so small models handle long inputs. It matched or beat long-context models on language modeling, question answering, and summarizing, with up to 57 times fewer parameters and less memory.
Abstract · HMT: Hierarchical Memory Transformer for Efficient Long Context Language Processing
Transformer-based large language models (LLM) have been widely used in language processing applications. However, due to the memory constraints of the devices, most of them restrict the context window. Even though recurrent models in previous works can memorize past tokens to enable unlimited context and maintain effectiveness, they have ``flat'' memory architectures. Such architectures have limitations in selecting and filtering information. Since humans are good at learning and self-adjustment, we believe that imitating brain memory hierarchy is beneficial for model memorization. Thus, we propose the Hierarchical Memory Transformer (HMT), a novel framework that facilitates a model's long-context processing ability by imitating human memorization behavior. Leveraging memory-augmented segment-level recurrence, we organize the memory hierarchy by preserving tokens from early input segments, passing memory embeddings along the sequence, and recalling relevant information from history. Evaluating general language modeling, question-answering tasks, and the summarization task, we show that HMT consistently improves the long-context processing ability of existing models. Furthermore, HMT achieves a comparable or superior generation quality to long-context LLMs with $2 \sim 57\times$ fewer parameters and $2.5 \sim 116\times$ less inference memory, significantly outperforming previous memory-augmented models. Code on Github: https://github.com/OswaldHe/HMT-pytorch.
Zifan He, Yingqi Cao, Zongyue Qin, Neha Prakriya, Yizhou Sun, Jason Cong
arXiv:2405.06067 · cs.CL, cs.LG · submitted May 9, 2024 · updated Feb 6, 2025
abstract · pdf · html · NAACL 2025 Main Conference
This looks really interesting. I've added the paper to my reading list and look forward to playing with the code. I'm curious to see what kinds of improvements we can get by agumenting Transformers and other generative sequence models with this and other mechanisms implementing hierarchical memory.[a]
Shouldn't the authors cite the work by Jeff Hawkins et al at Numenta? Hawkins has been proposing AI models with hierarchical temporal memory for a long time.[b] I can't help but wonder if there is a way, somehow, to incorporate his work and ideas in Transformers and other generative sequence models.
We sure live in interesting times!
---
[a] In the past, I've experimented with mechanisms that add memory to Transformers, but never with hierarchy.
[b] https://en.wikipedia.org/wiki/Hierarchical_temporal_memory