In plain words: The model writes the instruction that would have produced each web document, then keeps only the best pairs to train a stronger model. After two rounds, it beat every rival built on the same base model that wasn't trained on a bigger teacher's outputs.
Abstract · Self-Alignment with Instruction Backtranslation
We present a scalable method to build a high quality instruction following language model by automatically labelling human-written text with corresponding instructions. Our approach, named instruction backtranslation, starts with a language model finetuned on a small amount of seed data, and a given web corpus. The seed model is used to construct training examples by generating instruction prompts for web documents (self-augmentation), and then selecting high quality examples from among these candidates (self-curation). This data is then used to finetune a stronger model. Finetuning LLaMa on two iterations of our approach yields a model that outperforms all other LLaMa-based models on the Alpaca leaderboard not relying on distillation data, demonstrating highly effective self-alignment.
Xian Li, Ping Yu, Chunting Zhou, Timo Schick, Omer Levy, Luke Zettlemoyer, Jason Weston, Mike Lewis
arXiv:2308.06259 · cs.CL · submitted Aug 11, 2023 · updated Mar 12, 2024
abstract · pdf · html · ICLR2024 camera ready