In plain words: A model is pre-trained on sequences produced by random computer programs before it ever sees real data, so it learns general pattern-prediction. Fine-tuning after this warm-up then converged faster and generalized better than the usual training from scratch.
Abstract
We investigate the use of randomly generated data for the sake of pre-training a model. We justify this approach theoretically from the perspective of algorithmic complexity, building on recent research that shows that sequence models can be trained to approximate Solomonoff induction. We derive similar, but complementary theoretical results. We show empirically that synthetically generated data can be used to pre-train a model before the data is seen. We replicate earlier results that models trained this way show zero-shot in-context learning across a variety of datasets, and that this performance improves with scale. We extend earlier results to real-world data, and show that finetuning a model after pre-training offers faster convergence and better generalization.
Peter Bloem
arXiv:2506.20057 · cs.LG · submitted Jun 24, 2025
abstract · pdf · html
It’s good to compare various model sizes and evaluation tasks and random data generators. I just think the paper would more effectively prove its point if it could show models of same sizes which see this random data can learn better from evaluation data later on.
Could even take the initial checkpoint of the model before universal pretraining against the pretrained checkpoint. If the method works, the one that did UP will win.
Maybe I’m way off, I’ll admit I only skimmed it so far. Seems promising, just wishing for some controls.