about
Direct Nash Optimization: Teaching language models to self-improve (arxiv.org)
52 points by tosh on Apr 8, 2024 | hide | past | pdf | 11 comments on HN

In plain words: Instead of scoring each answer and maximizing it, this method compares answers in pairs, handling preferences that loop in circles, and improves the model each round. A 7-billion-parameter model trained this way beat GPT-4-Turbo 33% of the time, up from 7%, outperforming larger models.

Abstract · Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

This paper studies post-training large language models (LLMs) using preference feedback from a powerful oracle to help a model iteratively improve over itself. The typical approach for post-training LLMs involves Reinforcement Learning from Human Feedback (RLHF), which traditionally separates reward learning and subsequent policy optimization. However, such a reward maximization approach is limited by the nature of "point-wise" rewards (such as Bradley-Terry model), which fails to express complex intransitive or cyclic preference relations. While advances on RLHF show reward learning and policy optimization can be merged into a single contrastive objective for stability, they yet still remain tethered to the reward maximization framework. Recently, a new wave of research sidesteps the reward maximization presumptions in favor of directly optimizing over "pair-wise" or general preferences. In this paper, we introduce Direct Nash Optimization (DNO), a provable and scalable algorithm that marries the simplicity and stability of contrastive learning with theoretical generality from optimizing general preferences. Because DNO is a batched on-policy algorithm using a regression-based objective, its implementation is straightforward and efficient. Moreover, DNO enjoys monotonic improvement across iterations that help it improve even over a strong teacher (such as GPT-4). In our experiments, a resulting 7B parameter Orca-2.5 model aligned by DNO achieves the state-of-the-art win-rate against GPT-4-Turbo of 33% on AlpacaEval 2.0 (even after controlling for response length), an absolute gain of 26% (7% to 33%) over the initializing model. It outperforms models with far more parameters, including Mistral Large, Self-Rewarding LM (70B parameters), and older versions of GPT-4.

Corby Rosset, Ching-An Cheng, Arindam Mitra, Michael Santacroce, Ahmed Awadallah, Tengyang Xie
arXiv:2404.03715 · cs.LG, cs.AI, cs.CL · submitted Apr 4, 2024
abstract · pdf · html

add comment on HN

> 7B parameter Orca-2.5 model aligned by DNO achieves the state-of-the-art win-rate against GPT-4-Turbo of 33% on AlpacaEval 2.0 (even after controlling for response length), an absolute gain of 26% (7%→33%) over the initializing model. It outperforms models with far more parameters, including Mistral Large, Self-Rewarding LM (70B parameters), and older versions of GPT-4

edit: updated quote w/ more context

Still only 33% though. Which is impressive for a 7B model, but the student has not yet surpassed the teacher
author here; yea the goal is to try and get the student to surpass the teacher, but if it can't, this is the best way to get close. For these contrastive losses, our intuition is that the model isn't trying to emulate the teacher so much as learning from the 'delta' between itself and the teacher
If you only emulate the teacher can you ever become an master?
It works in many competitive fields
Surely it needs more than just emulating the teacher, though, it would require exploring beyond those limitations?
But eventually your human mentor diminishes with age.

And then you can crush them.

While the math seems intimidating, it does not look all too different from SPIN and previous researches. Pretty surprising how effective this is though. The costs here seems to be way higher too (with all the GPT4 calls).

> We also do a brief cost analysis associated with the scaled-up experiment on 600k training inputs. The major line items are the cost of sampling outputs, annotating them with GPT-4 to construct training pairs, and then training the next iteration against those pairs. For _each_ of the six iterations:

> 1. Sampling: it took about 18-24 hours to inference 5 outputs for all 100k examples on 10 8xA100 80GB pods, depending on the average length, costing about $6,000 based on spot pricing.

> 2. Annotation: the average number of prompt tokens sent to GPT-4 for annotation across iterations was about 450M, with an average of about 60M completion tokens, amounting to about $34,000 based on the version of the endpoint we were using.

> 3. Training: ironically, training was the cheapest step, taking only 12-24 hours on two 8xA100 80GB nodes

Yeah most of the cost in the future will be in preparing the training data by doing inference. This is the only way models can learn from their mistakes.
Do papers like these ever make the data they generate publically avalible or are we expected to pay the same api fees if we ever want to verify the work?
What is the UI for human preference then? Is it still just asking people to pick the best of two options?