In plain words: An already-trained big language model gets a few extra training steps where it practices filling in missing or scrambled text. This tiny add-on—about 0.1% more computing—let the largest version match the original's performance using roughly half the total training compute.
Abstract
Scaling language models improves performance but comes with significant computational costs. This paper proposes UL2R, a method that substantially improves existing language models and their scaling curves with a relatively tiny amount of extra compute. The key idea is to continue training a state-of-the-art large language model (e.g., PaLM) on a few more steps with UL2's mixture-of-denoiser objective. We show that, with almost negligible extra computational costs and no new sources of data, we are able to substantially improve the scaling properties of large language models on downstream metrics. In this paper, we continue training PaLM with UL2R, introducing a new set of models at 8B, 62B, and 540B scale which we call U-PaLM. Impressively, at 540B scale, we show an approximately 2x computational savings rate where U-PaLM achieves the same performance as the final PaLM 540B model at around half its computational budget (i.e., saving $\sim$4.4 million TPUv4 hours). We further show that this improved scaling curve leads to 'emergent abilities' on challenging BIG-Bench tasks -- for instance, U-PaLM does much better than PaLM on some tasks or demonstrates better quality at much smaller scale (62B as opposed to 540B). Overall, we show that U-PaLM outperforms PaLM on many few-shot setups, i.e., English NLP tasks (e.g., commonsense reasoning, question answering), reasoning tasks with chain-of-thought (e.g., GSM8K), multilingual tasks (MGSM, TydiQA), MMLU and challenging BIG-Bench tasks. Finally, we provide qualitative examples showing the new capabilities of U-PaLM for single and multi-span infilling.
Yi Tay, Jason Wei, Hyung Won Chung, Vinh Q. Tran, David R. So, Siamak Shakeri, Xavier Garcia, Huaixiu Steven Zheng, Jinfeng Rao, Aakanksha Chowdhery, Denny Zhou, Donald Metzler, et al.
arXiv:2210.11399 · cs.CL, cs.AI, cs.LG · submitted Oct 20, 2022 · updated Nov 16, 2022
abstract · pdf · html · V2 has updated references/related work
- Training on a mixture of fill-in-the-gaps (a few missing words) and denoising (every word slightly corrupted) produces better LLMs than either one alone.
- This advantage (that of using both metrics at once) can be gained with just a little extra training on a model previously trained with only one of them.
This results in 2-4% improvements on most tasks, with a couple really big improvements (+20%) and one quite surprising (+60%) on a few BigBench tasks. The large percent improvements on the BigBench tasks seem to have more to do with the low initial performance than the new performance being outstanding; the 60% was from 7.6% right to 12.5% right.