about
BTLM-3B-8K: 7B Parameter Performance in a 3B Parameter Model (arxiv.org)
3 points by jwan584 on Sep 22, 2023 | hide | past | pdf | 2 comments on HN

In plain words: A 3-billion-parameter language model was trained on cleaned, deduplicated text with tuned training settings and long-context support to reach the quality of models twice its size. It beat every other 3-billion model by 2–5.5% on downstream tasks while needing just 3GB of memory.

Abstract

We introduce the Bittensor Language Model, called "BTLM-3B-8K", a new state-of-the-art 3 billion parameter open-source language model. BTLM-3B-8K was trained on 627B tokens from the SlimPajama dataset with a mixture of 2,048 and 8,192 context lengths. BTLM-3B-8K outperforms all existing 3B parameter models by 2-5.5% across downstream tasks. BTLM-3B-8K is even competitive with some 7B parameter models. Additionally, BTLM-3B-8K provides excellent long context performance, outperforming MPT-7B-8K and XGen-7B-8K on tasks up to 8,192 context length. We trained the model on a cleaned and deduplicated SlimPajama dataset; aggressively tuned the \textmu P hyperparameters and schedule; used ALiBi position embeddings; and adopted the SwiGLU nonlinearity. On Hugging Face, the most popular models have 7B parameters, indicating that users prefer the quality-size ratio of 7B models. Compacting the 7B parameter model to one with 3B parameters, with little performance impact, is an important milestone. BTLM-3B-8K needs only 3GB of memory with 4-bit precision and takes 2.5x less inference compute than 7B models, helping to open up access to a powerful language model on mobile and edge devices. BTLM-3B-8K is available under an Apache 2.0 license on Hugging Face: https://huggingface.co/cerebras/btlm-3b-8k-base.

Nolan Dey, Daria Soboleva, Faisal Al-Khateeb, Bowen Yang, Ribhu Pathria, Hemant Khachane, Shaheer Muhammad, Zhiming, Chen, Robert Myers, Jacob Robert Steeves, Natalia Vassilieva, et al.
arXiv:2309.11568 · cs.AI, cs.CL, cs.LG · submitted Sep 20, 2023
abstract · pdf · html

add comment on HN
Also discussed: Sep 2023 (2 points, 0 comments)

When models are released like this, it would be great to do it with a PR to ggml/llama.cpp giving support, or use a format that's already supported. Imo if I'm choosing between a 3B and a 7B, I'm using it in an edge or local model and I don't want HF/pytorch. It would be easier to evaluate and rank higher in things to consider if I could easily get it into llama.cpp.
A helpful paper with the full recipe Cerebras uses to train LLMs and their process including: - Extensively deduplicated dataset (SlimPajama) - Hyperparameter search using muP - Variable sequence length training + ALiBi - Aggressive LR decay