about
Curious Decline of Linguistic Diversity: Training LLMs on Synthetic Text (2023) (arxiv.org)
2 points by 1vuio0pswjnm7 on Mar 7, 2024 | hide | past | pdf | 1 comment on HN

In plain words: They measured how varied the words, sentence shapes, and ideas are in text produced by models trained repeatedly on earlier models' writing. That variety shrank with every round, most sharply for tasks that call for creativity.

Abstract · The Curious Decline of Linguistic Diversity: Training Language Models on Synthetic Text

This study investigates the consequences of training language models on synthetic data generated by their predecessors, an increasingly prevalent practice given the prominence of powerful generative models. Diverging from the usual emphasis on performance metrics, we focus on the impact of this training methodology on linguistic diversity, especially when conducted recursively over time. To assess this, we adapt and develop a set of novel metrics targeting lexical, syntactic, and semantic diversity, applying them in recursive finetuning experiments across various natural language generation tasks in English. Our findings reveal a consistent decrease in the diversity of the model outputs through successive iterations, especially remarkable for tasks demanding high levels of creativity. This trend underscores the potential risks of training language models on synthetic text, particularly concerning the preservation of linguistic richness. Our study highlights the need for careful consideration of the long-term effects of such training approaches on the linguistic capabilities of language models.

Yanzhu Guo, Guokan Shang, Michalis Vazirgiannis, Chloé Clavel
arXiv:2311.09807 · cs.CL · submitted Nov 16, 2023 · updated Apr 16, 2024
abstract · pdf · html · Accepted to NAACL 2024 Findings

add comment on HN

> aimed at addressing the limited supply of human-generated training data.

Does that mean that we’ve reached a point where they’ve consumed all information that humans have created (and that we have in digital form) and yet they do not approach human intelligence and self-awareness and common sense?