In plain words: Instead of just answering and hoping, the model gives an answer, then rates how likely it is correct. Bigger models were well calibrated — their stated confidence matched how often they were right — and could also predict beforehand which questions they knew.
Abstract
We study whether language models can evaluate the validity of their own claims and predict which questions they will be able to answer correctly. We first show that larger models are well-calibrated on diverse multiple choice and true/false questions when they are provided in the right format. Thus we can approach self-evaluation on open-ended sampling tasks by asking models to first propose answers, and then to evaluate the probability "P(True)" that their answers are correct. We find encouraging performance, calibration, and scaling for P(True) on a diverse array of tasks. Performance at self-evaluation further improves when we allow models to consider many of their own samples before predicting the validity of one specific possibility. Next, we investigate whether models can be trained to predict "P(IK)", the probability that "I know" the answer to a question, without reference to any particular proposed answer. Models perform well at predicting P(IK) and partially generalize across tasks, though they struggle with calibration of P(IK) on new tasks. The predicted P(IK) probabilities also increase appropriately in the presence of relevant source materials in the context, and in the presence of hints towards the solution of mathematical word problems. We hope these observations lay the groundwork for training more honest models, and for investigating how honesty generalizes to cases where models are trained on objectives other than the imitation of human writing.
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, et al.
arXiv:2207.05221 · cs.CL, cs.AI, cs.LG · submitted Jul 11, 2022 · updated Nov 21, 2022
abstract · pdf · html · 23+17 pages; refs added, typos fixed
> With this fundraise, we’re going to explore the predictable scaling properties of machine learning systems, while closely examining the unpredictable ways in which capabilities and safety issues can emerge at scale
https://venturebeat.com/2022/06/27/10-new-ai-unicorns-flying...
Estimating confidence in predictions is an important safety issue. It's usually very hard to do correctly for small models. Fortunately large models have read everything so they have less unknown unknowns.