about
Small Language Models: Survey, Measurements, and Insights (arxiv.org)
1 point by sebg on Sep 26, 2024 | hide | past | pdf | discuss on HN

In plain words: A survey of 70 open-source small language models compares their designs, training data, and recipes, and tests reasoning, math, and long-context skills. Unlike surveys that only score accuracy, it also measures how fast each runs and how much memory it needs on a device.

Abstract

Small language models (SLMs), despite their widespread adoption in modern smart devices, have received significantly less academic attention compared to their large language model (LLM) counterparts, which are predominantly deployed in data centers and cloud environments. While researchers continue to improve the capabilities of LLMs in the pursuit of artificial general intelligence, SLM research aims to make machine intelligence more accessible, affordable, and efficient for everyday tasks. Focusing on transformer-based, decoder-only language models with 100M-5B parameters, we survey 70 state-of-the-art open-source SLMs, analyzing their technical innovations across three axes: architectures, training datasets, and training algorithms. In addition, we evaluate their capabilities in various domains, including commonsense reasoning, mathematics, in-context learning, and long context. To gain further insight into their on-device runtime costs, we benchmark their inference latency and memory footprints. Through in-depth analysis of our benchmarking data, we offer valuable insights to advance research in this field.

Zhenyan Lu, Xiang Li, Dongqi Cai, Rongjie Yi, Fangming Liu, Xiwen Zhang, Nicholas D. Lane, Mengwei Xu
arXiv:2409.15790 · cs.CL, cs.AI, cs.LG · submitted Sep 24, 2024 · updated Feb 26, 2025
abstract · pdf · html

add comment on HN
Also discussed: Sep 2024 (1 point, 0 comments)