about
The cringe Loss: Learning what language not to model (arxiv.org)
1 point by PaulHoule on Nov 14, 2022 | hide | past | pdf | 1 comment on HN

In plain words: Instead of training a chatbot only on good human text, this approach pairs each good response with a bad one and teaches the model to prefer the good and avoid the bad. It beat strong competitors on safe replies, staying consistent, and open-ended chat.

Abstract · The CRINGE Loss: Learning what language not to model

Standard language model training employs gold human documents or human-human interaction data, and treats all training data as positive examples. Growing evidence shows that even with very large amounts of positive training data, issues remain that can be alleviated with relatively small amounts of negative data -- examples of what the model should not do. In this work, we propose a novel procedure to train with such data called the CRINGE loss (ContRastive Iterative Negative GEneration). We show the effectiveness of this approach across three different experiments on the tasks of safe generation, contradiction avoidance, and open-domain dialogue. Our models outperform multiple strong baselines and are conceptually simple, easy to train and implement.

Leonard Adolphs, Tianyu Gao, Jing Xu, Kurt Shuster, Sainbayar Sukhbaatar, Jason Weston
arXiv:2211.05826 · cs.CL, cs.AI · submitted Nov 10, 2022
abstract · pdf · html

add comment on HN

Might not want to use fancy fonts in titles - "Please don't do things to make titles stand out"

https://news.ycombinator.com/newsguidelines.html