In plain words: They studied how AI-generated data mixed into training data hurts learning, using a simple prediction setup with math proofs and experiments. Even 1% fake data can stop bigger training sets from helping, and bigger models can make the damage worse.
Abstract · Strong Model Collapse
Within the scaling laws paradigm, which underpins the training of large neural networks like ChatGPT and Llama, we consider a supervised regression setting and establish the existance of a strong form of the model collapse phenomenon, a critical performance degradation due to synthetic data in the training corpus. Our results show that even the smallest fraction of synthetic data (e.g., as little as 1\% of the total training dataset) can still lead to model collapse: larger and larger training sets do not enhance performance. We further investigate whether increasing model size, an approach aligned with current trends in training large language models, exacerbates or mitigates model collapse. In a simplified regime where neural networks are approximated via random projections of tunable size, we both theoretically and empirically show that larger models can amplify model collapse. Interestingly, our theory also indicates that, beyond the interpolation threshold (which can be extremely high for very large datasets), larger models may mitigate the collapse, although they do not entirely prevent it. Our theoretical findings are empirically verified through experiments on language models and feed-forward neural networks for images.
Elvis Dohmatob, Yunzhen Feng, Arjun Subramonian, Julia Kempe
arXiv:2410.04840 · cs.LG, stat.ML · submitted Oct 7, 2024 · updated Oct 8, 2024
abstract · pdf · html
This paper that i read a few days ago establishes the existance of a strong model collapse, and says that (feom their own words "as little as 1\% of the trainig data") can still lead to model collapse.
What i found interesting beyond their results, which boil down to: Model collapse can be steep, only diminishing as the ration of synthetic data/real dat gets smaller, and that larger models experience a more severe model collapse. Is that in tge end they talk about the importance of labeling and curating real data.
How would real datavbe labelled? Wouldn't it be easier to labell generated data? And isnt the internet already stuffed with Generated images, SEO llm articles and more? All of which are becoming harder and harder to detect.
I dont want to call this a pre war steel situation since humans are still making human made content, but with llm made SE optimised content flooding the web on mass, amd the first three image results on google being AI generated, its starting to seem more like finding a needle in a haystack.
What do you guys think?