about
Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence (arxiv.org)
1 point by tosh on Apr 10, 2024 | hide | past | pdf | discuss on HN

In plain words: Eagle and Finch give a running-memory language model a richer matrix memory per head and input-dependent updates, so it still reads text cheaply one step at a time. Trained on 1.12 trillion multilingual tokens, they came close to the best models on many tests.

Abstract

We present Eagle (RWKV-5) and Finch (RWKV-6), sequence models improving upon the RWKV (RWKV-4) architecture. Our architectural design advancements include multi-headed matrix-valued states and a dynamic recurrence mechanism that improve expressivity while maintaining the inference efficiency characteristics of RNNs. We introduce a new multilingual corpus with 1.12 trillion tokens and a fast tokenizer based on greedy matching for enhanced multilinguality. We trained four Eagle models, ranging from 0.46 to 7.5 billion parameters, and two Finch models with 1.6 and 3.1 billion parameters and find that they achieve competitive performance across a wide variety of benchmarks. We release all our models on HuggingFace under the Apache 2.0 license. Models at: https://huggingface.co/RWKV Training code at: https://github.com/RWKV/RWKV-LM Inference code at: https://github.com/RWKV/ChatRWKV Time-parallel training code at: https://github.com/RWKV/RWKV-infctx-trainer

Bo Peng, Daniel Goldstein, Quentin Anthony, Alon Albalak, Eric Alcaide, Stella Biderman, Eugene Cheah, Xingjian Du, Teddy Ferdinan, Haowen Hou, Przemysław Kazienko, Kranthi Kiran GV, et al.
arXiv:2404.05892 · cs.CL, cs.AI · submitted Apr 8, 2024 · updated Sep 26, 2024
abstract · pdf · html

add comment on HN
Also discussed: Apr 2024 (1 point, 0 comments)