about
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models (arxiv.org)
41 points by ColinWright on Jan 3, 2024 | hide | past | pdf | 12 comments on HN

In plain words: A weak language model improves itself by writing answers, then learning to tell its own replies apart from human-written ones, needing no new human labels. Across several benchmarks it beat a rival training approach that adds extra AI-judged preference labels.

Abstract

Harnessing the power of human-annotated data through Supervised Fine-Tuning (SFT) is pivotal for advancing Large Language Models (LLMs). In this paper, we delve into the prospect of growing a strong LLM out of a weak one without the need for acquiring additional human-annotated data. We propose a new fine-tuning method called Self-Play fIne-tuNing (SPIN), which starts from a supervised fine-tuned model. At the heart of SPIN lies a self-play mechanism, where the LLM refines its capability by playing against instances of itself. More specifically, the LLM generates its own training data from its previous iterations, refining its policy by discerning these self-generated responses from those obtained from human-annotated data. Our method progressively elevates the LLM from a nascent model to a formidable one, unlocking the full potential of human-annotated demonstration data for SFT. Theoretically, we prove that the global optimum to the training objective function of our method is achieved only when the LLM policy aligns with the target data distribution. Empirically, we evaluate our method on several benchmark datasets including the HuggingFace Open LLM Leaderboard, MT-Bench, and datasets from Big-Bench. Our results show that SPIN can significantly improve the LLM's performance across a variety of benchmarks and even outperform models trained through direct preference optimization (DPO) supplemented with extra GPT-4 preference data. This sheds light on the promise of self-play, enabling the achievement of human-level performance in LLMs without the need for expert opponents. Codes are available at https://github.com/uclaml/SPIN.

Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, Quanquan Gu
arXiv:2401.01335 · cs.LG, cs.AI, cs.CL, stat.ML · submitted Jan 2, 2024 · updated Jun 14, 2024
abstract · pdf · html · 22 pages, 6 figures, 7 tables. In ICML 2024

add comment on HN

They say it doesn’t need preference data, but it seems to me that this does use preference data - the preferred response is from GPt-4, and the non-preferred response is from their model. It doesn’t fundamentally obviate the need to collect a high quality dataset from somewhere else.

In AlphaGo self play, the only external data was grandmaster Go moves that were used in a first pretraining phase of the policy network, and in AlphaGo Zero there was no external data at all. That’s what I would understand as self play really.

Seems to be more efficient than DPO - will try it out to compare

> It doesn’t fundamentally obviate the need to collect a high quality dataset from somewhere else.

You'd think this would be obvious: there's only so much worth you can "mine" out of any given data set. We've probably got lots of room left to further refine existing datasets, but clearly you need a high-quality dataset to start with to learn anything at all, right? And yet I wonder ...

> In AlphaGo Zero there was no external data at all. That’s what I would understand as self play really.

The interesting thing here is the complexity of the strategy -- the crazy amount there is to learn -- vs. the size of the ruleset. The raw materials being mined for strategy here is simply the simple rules, and yet they provide quite an impressive amount of possible learning. Similarly, there seems to be an endless amount that number theorists can learn from the simple rules of integer addition and multiplication.

I can't help but think about Chomsky, about the possibility of an implicit grammar, or psuedo-grammar, or grammar-factory, that all humans seem born with. Is there some threshold of examples of human language beyond which an LLM can inductively reason out the "rules" of our innate grammar-factory, and from there, self-play until it has mastered language as well as humans have? As well as humans ever could?

I can't really take seriously any research with "elevates the LLM from a nascent model to a formidable one" in the abstract...

If you want to catch my attention, say "+XX% at [benchmark], with the same number of weights and training data".

Just read the paper, or open the paper and scroll down to the first bar graph you see.
Tuning down these qualifiers perceived as subjective and focusing on the important stuff is a fine skill to develop. Maybe this is a useful research for you and you are passing it down because of cosmetics.
That's especially valuable since it's normally impossible to take even the benchmark percentage increase claims seriously for LLM research, because they're usually reached through training set contamination of the test questions. But your requirement of not introducing any new outside training data to an existing published model would take care of that.
In their defence, the authors are

> Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, Quanquan Gu

I'm guessing from how the introduction is written they aren't native English speakers, so they shouldn't be judged as if they are. Presumably the stuff they write sounds more normal in their native language(s).

The authors all have UCLA email addresses. Having a foreign name doesn't mean you can't speak English.
I never said that having a foreign name doesn't mean you can't speak English, but writing a serious paper where the first sentence starts

> Large Language Models (LLMs) have began a groundbreaking era in artificial general intelligence

Pretty strongly suggests that your work should be judged on its scientific (and not linguistic) merit.

I considered that... but no matter the language, in science, numbers are always preferred over "wow big amazing woo!"
what a chaotic way to name that backronym
This seems like a very clever idea because it is so obvious. Is this the kind of thing that OpenAI will be doing anyway behind closed doors?