about
The Capacity for Moral Self-Correction in Large Language Models (arxiv.org)
3 points by georgehill on Feb 16, 2023 | hide | past | pdf | discuss on HN

In plain words: They tested whether language models trained on human ratings can simply be told to avoid harmful output and actually do it. The ability to self-correct appeared at 22 billion parameters and got better as models grew larger and received more human-feedback training.

Abstract

We test the hypothesis that language models trained with reinforcement learning from human feedback (RLHF) have the capability to "morally self-correct" -- to avoid producing harmful outputs -- if instructed to do so. We find strong evidence in support of this hypothesis across three different experiments, each of which reveal different facets of moral self-correction. We find that the capability for moral self-correction emerges at 22B model parameters, and typically improves with increasing model size and RLHF training. We believe that at this level of scale, language models obtain two capabilities that they can use for moral self-correction: (1) they can follow instructions and (2) they can learn complex normative concepts of harm like stereotyping, bias, and discrimination. As such, they can follow instructions to avoid certain kinds of morally harmful outputs. We believe our results are cause for cautious optimism regarding the ability to train language models to abide by ethical principles.

Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas I. Liao, Kamilė Lukošiūtė, Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, Dawn Drain, Dustin Li, et al.
arXiv:2302.07459 · cs.CL · submitted Feb 15, 2023 · updated Feb 18, 2023
abstract · pdf · html

add comment on HN