about
RecurrentGemma: Moving Past Transformers for Efficient Open Language Models (arxiv.org)
47 points by CharlesW on Apr 22, 2024 | hide | past | pdf | 3 comments on HN

In plain words: Instead of letting every word attend to every word, this language model keeps a fixed-size running memory and only looks back a short way, so memory stays flat even on long texts. It matches similar-size transformer-based Gemma models despite training on fewer words.

Abstract

We introduce RecurrentGemma, a family of open language models which uses Google's novel Griffin architecture. Griffin combines linear recurrences with local attention to achieve excellent performance on language. It has a fixed-sized state, which reduces memory use and enables efficient inference on long sequences. We provide two sizes of models, containing 2B and 9B parameters, and provide pre-trained and instruction tuned variants for both. Our models achieve comparable performance to similarly-sized Gemma baselines despite being trained on fewer tokens.

Aleksandar Botev, Soham De, Samuel L Smith, Anushan Fernando, George-Cristian Muraru, Ruba Haroun, Leonard Berrada, Razvan Pascanu, Pier Giuseppe Sessa, Robert Dadashi, Léonard Hussenot, Johan Ferret, et al.
arXiv:2404.07839 · cs.LG, cs.AI, cs.CL · submitted Apr 11, 2024 · updated Aug 28, 2024
abstract · pdf · html

add comment on HN

I'd be curious to see this architecture trained to a comparable level to Llama3.

8B params, 15T tokens of training.

Llama3 8B supercedes ChatGPT3.5. If RecurrentGemma scales, then it would be both faster and more capable than its Llama peer.