about
LLaMA Pro: Progressive LLaMA with Block Expansion (arxiv.org)
2 points by DarmokJalad1701 on Jan 5, 2024 | hide | past | pdf | 1 comment on HN

In plain words: New blocks are bolted onto a trained language model and only those blocks are taught code and math text, so old knowledge stays intact. The grown model, from a 7-billion-parameter one, beats other open models in its family on general, coding, and math tests.

Abstract

Humans generally acquire new skills without compromising the old; however, the opposite holds for Large Language Models (LLMs), e.g., from LLaMA to CodeLLaMA. To this end, we propose a new post-pretraining method for LLMs with an expansion of Transformer blocks. We tune the expanded blocks using only new corpus, efficiently and effectively improving the model's knowledge without catastrophic forgetting. In this paper, we experiment on the corpus of code and math, yielding LLaMA Pro-8.3B, a versatile foundation model initialized from LLaMA2-7B, excelling in general tasks, programming, and mathematics. LLaMA Pro and its instruction-following counterpart (LLaMA Pro-Instruct) achieve advanced performance among various benchmarks, demonstrating superiority over existing open models in the LLaMA family and the immense potential of reasoning and addressing diverse tasks as an intelligent agent. Our findings provide valuable insights into integrating natural and programming languages, laying a solid foundation for developing advanced language agents that operate effectively in various environments.

Chengyue Wu, Yukang Gan, Yixiao Ge, Zeyu Lu, Jiahao Wang, Ye Feng, Ying Shan, Ping Luo
arXiv:2401.02415 · cs.CL · submitted Jan 4, 2024 · updated May 30, 2024
abstract · pdf · html · Accepted by ACL 2024, Main Conference

add comment on HN

Appears to be a way to address catastrophic forgetting in LLMs