about
Learning to summarize from human feedback (2022) (arxiv.org)
56 points by georgehill on Mar 4, 2023 | hide | past | pdf | 12 comments on HN

In plain words: Instead of copying human-written summaries, the system learns from people picking which of two summaries they prefer, then uses that preference predictor as a score to improve the summarizer. People preferred its summaries over human-written ones and over bigger models trained the usual way.

Abstract · Learning to summarize from human feedback

As language models become more powerful, training and evaluation are increasingly bottlenecked by the data and metrics used for a particular task. For example, summarization models are often trained to predict human reference summaries and evaluated using ROUGE, but both of these metrics are rough proxies for what we really care about -- summary quality. In this work, we show that it is possible to significantly improve summary quality by training a model to optimize for human preferences. We collect a large, high-quality dataset of human comparisons between summaries, train a model to predict the human-preferred summary, and use that model as a reward function to fine-tune a summarization policy using reinforcement learning. We apply our method to a version of the TL;DR dataset of Reddit posts and find that our models significantly outperform both human reference summaries and much larger models fine-tuned with supervised learning alone. Our models also transfer to CNN/DM news articles, producing summaries nearly as good as the human reference without any news-specific fine-tuning. We conduct extensive analyses to understand our human feedback dataset and fine-tuned models We establish that our reward model generalizes to new datasets, and that optimizing our reward model results in better summaries than optimizing ROUGE according to humans. We hope the evidence from our paper motivates machine learning researchers to pay closer attention to how their training loss affects the model behavior they actually want.

Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, Paul Christiano
arXiv:2009.01325 · cs.CL, cs.AI, cs.LG · submitted Sep 2, 2020 · updated Feb 15, 2022
abstract · pdf · html · NeurIPS 2020

add comment on HN
Also discussed: Sep 2020 (2 points, 0 comments)

Note that this is pretty old (2020).

They released code, models and raw data here: https://github.com/openai/summarize-from-feedback

Feel like automating the human feedback, not the summaries themselves, that should have been the core focus of research like this. As is, even reading guidelines for summary evaluation they presented the reviewers are not reproducible.
Sorry if the answer is obvious but can we use this for our own usage? If yes, how?
I'd be curious to see how this does compared to models trained on more professional datasets than reddit tldr.

For example, train a model(s) by reading every single article (including paywall/cache replacement) of https://www.techmeme.com/river https://www.mediagazer.com/river https://www.memeorandum.com/river https://www.wesmirch.com/river https://ballbug.com/river and comparing it to the summary headline.

https://github.com/CurationCorp/curation-corpus

There's a dataset that would help there.

I’d still be curious how that compares to techmeme’s headline rewriting since 2013/2015.
What are these websites??
https://www.cnbc.com/amp/2017/03/22/meet-the-man-whose-site-...

https://www.businessinsider.com/techmeme-growth-2014-3

The river is the reverse chronological order similar to https://hckrnews.com

If you go back to the main page of tech meme, and hover over a story an arrow will appear on the left. Click it to see follow up stories to the first story.

The first two techmeme and mediagazer at least, do headline rewriting to debuzzfeed/deupworthy clickbait headlines. https://finance.yahoo.com/news/aggregators-attack-techmeme-h... https://www.poynter.org/reporting-editing/2015/techmeme-is-p...

Does anyone have a TL;DR on this?
I mean papers come with abstracts, but yeah:

> Open AI

> We collect a large, high-quality dataset of human comparisons between summaries, train a model to predict the human-preferred summary, and use that model as a reward function to fine-tune a summarization policy using reinforcement learning. We apply our method to a version of the TL;DR dataset of Reddit posts and find that our models significantly outperform both human reference summaries and much larger models fine-tuned with supervised learning alone.

tl;dr^2 they did ChatGPT to summaries

https://openai.com/blog/chatgpt#methods

Your tl;dr shows how text+reader_context can generate the best summaries. Those 5 words are perfect if you know what they did for ChatGPT.

This makes me think that to get high quality summaries: 1) they have to be generated for each individual reader 2) the AI should know what the reader knows

You achieved 2 by imagining what HN people may already know, but the ultimate goal would be to know what the individual reader knows.

And this chain of thought leads to - Perhaps all AI output should be generated on the fly for the end user with full (relevant and compressed) context. - Giving AI what we know is extremely dangerous if someone wants to use it for something bad (so we really, really need local AI).

Sure, these are all things people in the industry already know, but there must be a lot of people like me who are just now thinking about it.

A note: "^2" added humor to make the reading light with very few bytes. Perhaps the machine would do this too, if the reader wants.

Only few things make me very very excited and sometimes very very sad like some AI developments. But the exciting part wins for now!

Yeah, I think context is really important for future summarizers. If you think of it like information theory, the goal is what information to convey _for the message to be successfully decoded_. So the amount of information needed is not universal but context-dependent on the decoder (here the reader).

Being able to generate summaries at various levels of depth would be a great efficiency to consuming a lot of content. No more skimming through articles and books written like the author was paid by the character. But if you want more depth, it’s there. Like an LoD slider for information