about
A (sorta) recent paper about model collapse has got me thinking (arxiv.org)
3 points by Wheatman on Oct 19, 2024 | hide | past | pdf | 3 comments on HN

In plain words: They studied how AI-generated data mixed into training data hurts learning, using a simple prediction setup with math proofs and experiments. Even 1% fake data can stop bigger training sets from helping, and bigger models can make the damage worse.

Abstract · Strong Model Collapse

Within the scaling laws paradigm, which underpins the training of large neural networks like ChatGPT and Llama, we consider a supervised regression setting and establish the existance of a strong form of the model collapse phenomenon, a critical performance degradation due to synthetic data in the training corpus. Our results show that even the smallest fraction of synthetic data (e.g., as little as 1\% of the total training dataset) can still lead to model collapse: larger and larger training sets do not enhance performance. We further investigate whether increasing model size, an approach aligned with current trends in training large language models, exacerbates or mitigates model collapse. In a simplified regime where neural networks are approximated via random projections of tunable size, we both theoretically and empirically show that larger models can amplify model collapse. Interestingly, our theory also indicates that, beyond the interpolation threshold (which can be extremely high for very large datasets), larger models may mitigate the collapse, although they do not entirely prevent it. Our theoretical findings are empirically verified through experiments on language models and feed-forward neural networks for images.

Elvis Dohmatob, Yunzhen Feng, Arjun Subramonian, Julia Kempe
arXiv:2410.04840 · cs.LG, stat.ML · submitted Oct 7, 2024 · updated Oct 8, 2024
abstract · pdf · html

add comment on HN
Also discussed: Oct 2024 (1 point, 2 comments)

Hi new guy here.

Im the kind of deranged lunatic that reads Arxiv papers for fun, and about a week ago i landed on this paper.

It says that larger models suffer from model collapse from its own data much harder, and from a lower percentage of synthetic data after a certain point.

This got me curious, with how many AI generated images(which some claim to be billions) and text. How would data scrapers be able to avoid indirectly training newer and larger models on their own data? I doubt they personally curate each line of text they train them on? So do they just ignore it?

Take a look at this thing https://news.ycombinator.com/newsguidelines.html about titles - random arxiv papers you found interesting along with commentary on why you thought it was interesting are fine things to post, you just can't put the commentary in your post's title.
Sorry, thanks for telling me.