about
Meta: Memory Layers at Scale (arxiv.org)
4 points by georgehill on Jan 2, 2025 | hide | past | pdf | discuss on HN

In plain words: Instead of doing more math per word, the model adds a lookup table that stores facts and pulls out the few entries it needs, adding knowledge without extra computation. It beat dense models using over twice the computing power, especially on factual questions.

Abstract · Memory Layers at Scale

Memory layers use a trainable key-value lookup mechanism to add extra parameters to a model without increasing FLOPs. Conceptually, sparsely activated memory layers complement compute-heavy dense feed-forward layers, providing dedicated capacity to store and retrieve information cheaply. This work takes memory layers beyond proof-of-concept, proving their utility at contemporary scale. On downstream tasks, language models augmented with our improved memory layer outperform dense models with more than twice the computation budget, as well as mixture-of-expert models when matched for both compute and parameters. We find gains are especially pronounced for factual tasks. We provide a fully parallelizable memory layer implementation, demonstrating scaling laws with up to 128B memory parameters, pretrained to 1 trillion tokens, comparing to base models with up to 8B parameters.

Vincent-Pierre Berges, Barlas Oğuz, Daniel Haziza, Wen-tau Yih, Luke Zettlemoyer, Gargi Ghosh
arXiv:2412.09764 · cs.CL, cs.AI · submitted Dec 12, 2024 · updated Dec 20, 2024
abstract · pdf · html

add comment on HN
Also discussed: Jan 2025 (4 points, 0 comments)