about
Emergent Abilities of Large Language Models (2022) (arxiv.org)
1 point by squircle on Oct 3, 2024 | hide | past | pdf | discuss on HN

In plain words: Some skills vanish in small language models and switch on only once models get big, like a light flipping on. Since these skills appear suddenly, charting small models' steady gains can't predict them, so bigger models may keep unlocking new surprises.

Abstract · Emergent Abilities of Large Language Models

Scaling up language models has been shown to predictably improve performance and sample efficiency on a wide range of downstream tasks. This paper instead discusses an unpredictable phenomenon that we refer to as emergent abilities of large language models. We consider an ability to be emergent if it is not present in smaller models but is present in larger models. Thus, emergent abilities cannot be predicted simply by extrapolating the performance of smaller models. The existence of such emergence implies that additional scaling could further expand the range of capabilities of language models.

Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, et al.
arXiv:2206.07682 · cs.CL · submitted Jun 15, 2022 · updated Oct 26, 2022
abstract · pdf · html · Transactions on Machine Learning Research (TMLR), 2022

add comment on HN
Also discussed: Feb 2026 (1 point, 0 comments) · May 2024 (2 points, 0 comments) · Apr 2023 (4 points, 1 comment) · Apr 2023 (3 points, 0 comments) · Feb 2023 (2 points, 1 comment)