In plain words: A compact language model with 1.1 billion settings was trained on about 1 trillion chunks of text, three passes over the data, reusing a popular large model's design plus speed tricks to cut cost. It beat free models of similar size on many tasks.
Abstract
We present TinyLlama, a compact 1.1B language model pretrained on around 1 trillion tokens for approximately 3 epochs. Building on the architecture and tokenizer of Llama 2, TinyLlama leverages various advances contributed by the open-source community (e.g., FlashAttention and Lit-GPT), achieving better computational efficiency. Despite its relatively small size, TinyLlama demonstrates remarkable performance in a series of downstream tasks. It significantly outperforms existing open-source language models with comparable sizes. Our model checkpoints and code are publicly available on GitHub at https://github.com/jzhang38/TinyLlama.
Peiyuan Zhang, Guangtao Zeng, Tianduo Wang, Wei Lu
arXiv:2401.02385 · cs.CL, cs.AI · submitted Jan 4, 2024 · updated Jun 4, 2024
abstract · pdf · html · Technical Report
But they did move down and that's what's important.
There should probably be more aggressive learning rate annealing for models trying to be Chinchilla-optimal instead of just cosine-with-warmup like every other model nowadays.