about
Do Language Models Plan Ahead for Future Tokens? (2024) (arxiv.org)
2 points by Alifatisk 76 days ago | hide | past | pdf | discuss on HN

In plain words: Language models stash information at one word that later words use. Training without rewarding help to future steps tested why: synthetic data showed deliberate stashing, but language modeling leaned toward the same features helping now and later, with stashing growing as models scaled. Hmm, that's 2 sentences but total words: sentence 1 = 20, sentence 2 = "synthetic data showed deliberate stashing, but language modeling leaned toward the same features helping now and later, with stashing growing as models scaled" = 23. Total 43. Good. Wait, but the colon makes it read as one flowing idea. Fine. Check: "Training without rewarding help to future steps tested why" — a bit compressed but understandable. Maybe "Training models without rewarding help to future steps tested why" = 22, total 45. Keep the shorter version. Actually let me reconsider clarity: "Training without rewarding help to future steps tested why" — the subject "Training" tested why. Slightly awkward. Alternative: "To test why, models were trained with no reward for helping future steps." = To1 test2 why3 models4 were5 trained6 with7 no8 reward9 for10 helping11 future12 steps13 = 13. S1 = 23. Total 46. Over. "Models trained with no reward for helping future steps tested why" = 11 words → S1 = 21, total 44. Good and clearer. Final: "Language models stash information at one word that later words use. Models trained with no reward for helping future steps tested why: synthetic data showed deliberate stashing, but language modeling leaned toward the same features helping now and later, with stashing growing as models scaled." Count total: S1 part: Language1 models2 stash3 information4 at5 one6 word7 that8 later9 words10 use11 Models12 trained13 with14 no15 reward16 for17 helping18 future19 steps20 tested21 why22 = 22. S2: synthetic1 data2 showed3 deliberate4 stashing5 but6 language7 modeling8 leaned9 toward10 the11 same12 features13 helping14 now15 and16 later17 with18 stashing19 growing20 as21 models22 scaled23 = 23. Total 45. Over by 1! Trim: "with stashing growing as models scaled" → "and stashing grew with model size" = 6 words, same. → "with stashing growing as models grew" = 6. Same. Drop "deliberate": "synthetic data showed clear stashing" — but then contrast lost. Alternatively drop "the same" → "same features" ungrammatical. Drop "language modeling leaned toward" → "language modeling favored" = saves 2 words: but1 language2 modeling3 favored4 the5 same6 features7 helping8 now9 and10 later11 = 11 vs 17. Saves 6 → total 39. But "favored" is a bit strong vs "more suggestive". Hmm, "leaned toward" better conveys "suggestive". Alternative trim: "Models trained with no reward for helping future steps tested why" → "Models with no reward for helping future steps tested why" = Models1 with2 no3 reward4 for5 helping6 future7 steps8 tested9 why10 = 10 → S1 = 21, total 44. But "Models with no reward" is slightly odd; "Models trained with no reward" clearer. Hmm, "Models trained without rewarding help to future steps tested why" = Models1 trained2 without3 rewarding4 help5 to6 future7 steps8 tested9 why10 = 10 → S1 = 21, total 44. Good and clear. Final text: "Language models stash information at one word that later words use. Models trained without rewarding help to future steps tested why: synthetic data showed deliberate stashing, but language modeling leaned toward the same features helping now and later, with stashing growing as models scaled." Total 44. Good. Two sentences. No jargon. Starts with "Language models" — fine. One check: "the same features helping now and later" — describes breadcrumbs. Good. "deliberate stashing" describes pre-caching. Good. "stashing growing as models scaled" — pre-caching increases with model scale. ✓ Output plain text.Language models stash information at one word that later words use. Models trained without rewarding help to future steps tested why: synthetic data showed deliberate stashing, but language modeling leaned toward the same features helping now and later, with stashing growing as models scaled.

Abstract · Do language models plan ahead for future tokens?

Do transformers "think ahead" during inference at a given position? It is known transformers prepare information in the hidden states of the forward pass at time step $t$ that is then used in future forward passes $t+τ$. We posit two explanations for this phenomenon: pre-caching, in which off-diagonal gradient terms present during training result in the model computing features at $t$ irrelevant to the present inference task but useful for the future, and breadcrumbs, in which features most relevant to time step $t$ are already the same as those that would most benefit inference at time $t+τ$. We test these hypotheses by training language models without propagating gradients to past timesteps, a scheme we formalize as myopic training. In a constructed synthetic data setting, we find clear evidence for pre-caching. In the autoregressive language modeling setting, our experiments are more suggestive of the breadcrumbs hypothesis, though pre-caching increases with model scale.

Wilson Wu, John X. Morris, Lionel Levine
arXiv:2404.00859 · cs.LG, cs.CL · submitted Apr 1, 2024 · updated Aug 1, 2024
abstract · pdf · html · 24 pages, 11 figures. Camera-ready for COLM 2024

add comment on HN
Also discussed: Apr 2024 (3 points, 0 comments)