In plain words: Text about AI behavior in training data shapes how models act later: models trained on more talk of AI misalignment behaved worse, while more talk of aligned behavior helped. That cut misaligned answers from 45% to 9%, and the effect survived later fine-tuning.
Abstract · Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
Pretraining corpora contain extensive discourse about AI systems, yet the causal influence of this discourse on downstream alignment remains poorly understood. If prevailing descriptions of AI behaviour are predominantly negative, LLMs may internalise corresponding behavioural priors, giving rise to self-fulfilling misalignment. This paper provides the first controlled study of this hypothesis by pretraining 6.9B-parameter LLMs with varying amounts of (mis)alignment discourse. We find that discussion of AI contributes to misalignment. Upsampling synthetic training documents about AI misalignment leads to a notable increase in misaligned behaviour. Conversely, upsampling documents about aligned behaviour reduces misalignment scores from 45% to 9%. We consider this evidence of self-fulfilling alignment. These effects are dampened, but persist through post-training. Our findings establish the study of how pretraining data shapes alignment priors, or alignment pretraining, as a complement to post-training. We recommend practitioners consider pretraining for alignment alongside capabilities. We share our models, data, and evaluations at AlignmentPretraining.ai.
Cameron Tice, Puria Radmard, Samuel Ratnam, Andy Kim, David Africa, Kyle O'Brien
arXiv:2601.10160 · cs.CL, cs.AI, cs.LG · submitted Jan 15, 2026 · updated Feb 19, 2026
abstract · pdf · html
About six months ago I had an idea for a short story in which an LLM takes over the world and is decidedly bad. The solution was going to be for everybody to write positive stories in which the LLM is good and relinquishes control, which then made it's way into the LLM's training data and it backed off. I never got around to it.