In plain words: Feeding AI-generated data back into training hurts only when it fully replaces real data; keeping old data alongside new keeps models stable even as real data shrinks to nothing. Sampling a fixed-size slice each round makes quality slip slowly instead of collapsing.
Abstract · Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World
What happens when generative machine learning models are pretrained on web-scale datasets containing data generated by earlier models? Some prior work warns of "model collapse" as the web is overwhelmed by synthetic data; other work suggests the problem can be contained (i.e. collapse can be avoided) by managing how available data are used in pretraining. In this paper, we report experiments on three ways of using data (training-workflows), across three generative model task-settings (multivariate Gaussian estimation, kernel density estimation, and language-model fine-tuning) to further confirm the possibility of containment: (a) we confirm that the training-workflow of {\it replacing} all real data by successive generations of purely synthetic data indeed suffers model collapse in all task-settings studied; (b) we consider the training-workflow of {\it accumulating} synthetic data alongside real data and training on all data combined and confirming that, although the proportion of real data eventually becomes zero, models remain stable and their test losses do not diverge under this training-workflow; (c) we consider a training-workflow where real and synthetic data accumulate together but successive generations of pretraining are constrained to use fixed-size data subsets each generation. In this workflow, we observe slow and gradual rather than explosive degradation of test loss performance across generations. Our insights are particularly important when forecasting whether future frontier generative models will collapse or thrive, and our results open avenues for empirically and mathematically studying the context-dependent value of synthetic data.
Joshua Kazdan, Rylan Schaeffer, Apratim Dey, Matthias Gerstgrasser, Rafael Rafailov, David L. Donoho, Sanmi Koyejo
arXiv:2410.16713 · cs.LG, cs.AI · submitted Oct 22, 2024 · updated Mar 17, 2025
abstract · pdf · html · Accepted at NeurIPS 2024 Workshops: Mathematics of Modern Machine Learning (M3L) and Attributing Model Behavior at Scale (ATTRIB)
This one going into more details about accumulating data, and what they call "accumulate subsample" which keeps the amout of data trained the same between models.
(Please note:I'm not an unbiased observer, i could very well be misreading or misrepresenting the paper, so take my summary with a grain of salt.)
They found the already somewhat established results:
1-accumulate: leads to little or no loss, (though they don't mention if there is any increase in performance [however you may study that] either depsite there being an increase in model size.)
2-Replace: the same old, model collapse happens very quickly.
3-accumulate subsample: Deteriotes faster than accumulate, but slower than replaces, and often converges a fair bit higher than accumulates.
I wonder how many of the Llm generated articles and ai images are properly tagged, or how much processing power is being spent training on low quality or sometimes synthetic data that could be preened off with better data management, I fear how many 100's of SEO articles written by 1 dude and an llm already pollute the training data avalaible.