about
StableLM 1.6B Technical Report – includes all data, training, strategy (arxiv.org)
1 point by omnipotent_i on Feb 29, 2024 | hide | past | pdf | 1 comment on HN

In plain words: A compact language model with 1.6 billion weights, trained on multilingual text and also tuned to follow instructions, with free downloadable versions. It beat every other openly available model under 2 billion weights at release, and is small enough for everyday devices.

Abstract · Stable LM 2 1.6B Technical Report

We introduce StableLM 2 1.6B, the first in a new generation of our language model series. In this technical report, we present in detail the data and training procedure leading to the base and instruction-tuned versions of StableLM 2 1.6B. The weights for both models are available via Hugging Face for anyone to download and use. The report contains thorough evaluations of these models, including zero- and few-shot benchmarks, multilingual benchmarks, and the MT benchmark focusing on multi-turn dialogues. At the time of publishing this report, StableLM 2 1.6B was the state-of-the-art open model under 2B parameters by a significant margin. Given its appealing small size, we also provide throughput measurements on a number of edge devices. In addition, we open source several quantized checkpoints and provide their performance metrics compared to the original model.

Marco Bellagente, Jonathan Tow, Dakota Mahan, Duy Phung, Maksym Zhuravinskyi, Reshinth Adithyan, James Baicoianu, Ben Brooks, Nathan Cooper, Ashish Datta, Meng Lee, Emad Mostaque, et al.
arXiv:2402.17834 · cs.CL, stat.ML · submitted Feb 27, 2024
abstract · pdf · html · 23 pages, 6 figures

add comment on HN

StableLM 1.6B is a strong multilingual LM trained for 2T tokens of multilingual dataset. It offers strong performance compared to other models of the same size range including Google’s Gemma, Microsoft’s Phi2 and ByteDance’s Qwen. The technical report discloses all of the necessary details to reproduce the training procedure including the data mix. Training details: RoPE, LayerNorm, No MLP bias terms, but retains for QKV terms. Global Batch Size - 8M tokens MFU during training - 54.5% Tokenizer - based out of OpenAI's tiktoken's CL100k_base with vocab size ~ 100k. Paper shows a new hybrid LR scheduler used while training the model. report: https://arxiv.org/abs/2402.17834 base model: https://huggingface.co/stabilityai/stablelm-2-1_6b zephyr - https://huggingface.co/stabilityai/stablelm-2-zephyr-1_6b Ollama - `ollama run stablelm2`