about
Language Models (Mostly) Know What They Know (arxiv.org)
43 points by PaulHoule on Jul 13, 2022 | hide | past | pdf | 6 comments on HN

In plain words: Instead of just answering and hoping, the model gives an answer, then rates how likely it is correct. Bigger models were well calibrated — their stated confidence matched how often they were right — and could also predict beforehand which questions they knew.

Abstract

We study whether language models can evaluate the validity of their own claims and predict which questions they will be able to answer correctly. We first show that larger models are well-calibrated on diverse multiple choice and true/false questions when they are provided in the right format. Thus we can approach self-evaluation on open-ended sampling tasks by asking models to first propose answers, and then to evaluate the probability "P(True)" that their answers are correct. We find encouraging performance, calibration, and scaling for P(True) on a diverse array of tasks. Performance at self-evaluation further improves when we allow models to consider many of their own samples before predicting the validity of one specific possibility. Next, we investigate whether models can be trained to predict "P(IK)", the probability that "I know" the answer to a question, without reference to any particular proposed answer. Models perform well at predicting P(IK) and partially generalize across tasks, though they struggle with calibration of P(IK) on new tasks. The predicted P(IK) probabilities also increase appropriately in the presence of relevant source materials in the context, and in the presence of hints towards the solution of mathematical word problems. We hope these observations lay the groundwork for training more honest models, and for investigating how honesty generalizes to cases where models are trained on objectives other than the imitation of human writing.

Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, et al.
arXiv:2207.05221 · cs.CL, cs.AI, cs.LG · submitted Jul 11, 2022 · updated Nov 21, 2022
abstract · pdf · html · 23+17 pages; refs added, typos fixed

add comment on HN

Authors from Anthropic, a splinter group from OpenAI

> With this fundraise, we’re going to explore the predictable scaling properties of machine learning systems, while closely examining the unpredictable ways in which capabilities and safety issues can emerge at scale

https://venturebeat.com/2022/06/27/10-new-ai-unicorns-flying...

Estimating confidence in predictions is an important safety issue. It's usually very hard to do correctly for small models. Fortunately large models have read everything so they have less unknown unknowns.

What I’ve found for simpler models is that probability calibration is the difference between a model you can use and ‘nothing more to see folks, please move on…’

I’ve found that calibration frequently isn’t that hard to do, but mostly people don’t do it (e.g. they announce that their model is 99% accurate for a disease that occurs 1 in 10,000 to which the answer is ‘you can beat that accuracy by saying nobody has the disease) or maybe the model sucks (e.g. a calibrated full text retrieval system will never claim its results are better than 70% likely to be relevant.)

I was expecting the proposed approach to be a prompt hack that gets the model to output a well-calibrated confidence value as text.

"How confident are you in your answer? Let's calibrate carefully."

While I appreciate findings like this it feels more like alchemy than science. "Oh, we found this prompt X and it works better than prompt Y but pretty likely there is at least one prompt Z which we don't know yet that works even better". In addition to this, benchmark evaluation data sets are cool and everything but they represent real world LM application environments to a very limited extent only.
> it feels more like alchemy than science.

I’d say it feels more like science than engineering to capture the same idea. The whole point is that we don’t know how it all works yet, hence the science.