about
SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales (arxiv.org)
28 points by jasondavies on Jun 4, 2024 | hide | past | pdf | 9 comments on HN

In plain words: A training setup teaches a language model to give a precise confidence score and a note on what it is unsure about, by summarizing where its reasoning chains disagree. Compared with just asking for confidence, the scores matched real accuracy better without hurting performance.

Abstract

Large language models (LLMs) often generate inaccurate or fabricated information and generally fail to indicate their confidence, which limits their broader applications. Previous work elicits confidence from LLMs by direct or self-consistency prompting, or constructing specific datasets for supervised finetuning. The prompting-based approaches have inferior performance, and the training-based approaches are limited to binary or inaccurate group-level confidence estimates. In this work, we present the advanced SaySelf, a training framework that teaches LLMs to express more accurate fine-grained confidence estimates. In addition, beyond the confidence scores, SaySelf initiates the process of directing LLMs to produce self-reflective rationales that clearly identify gaps in their parametric knowledge and explain their uncertainty. This is achieved by using an LLM to automatically summarize the uncertainties in specific knowledge via natural language. The summarization is based on the analysis of the inconsistency in multiple sampled reasoning chains, and the resulting data is utilized for supervised fine-tuning. Moreover, we utilize reinforcement learning with a meticulously crafted reward function to calibrate the confidence estimates, motivating LLMs to deliver accurate, high-confidence predictions and to penalize overconfidence in erroneous outputs. Experimental results in both in-distribution and out-of-distribution datasets demonstrate the effectiveness of SaySelf in reducing the confidence calibration error and maintaining the task performance. We show that the generated self-reflective rationales are reasonable and can further contribute to the calibration. The code is made public at https://github.com/xu1868/SaySelf.

Tianyang Xu, Shujin Wu, Shizhe Diao, Xiaoze Liu, Xingyao Wang, Yangyi Chen, Jing Gao
arXiv:2405.20974 · cs.CL, cs.AI, cs.LG · submitted May 31, 2024 · updated Oct 4, 2024
abstract · pdf · html · EMNLP 2024 Main

add comment on HN

Excellent background on knowledge calibration from Anthropic:

https://arxiv.org/abs/2207.05221

"Calibration" in a knowledge context means having estimated_p(correct) ~ p(correct), and it turns out that LLMs are reasonably good at this. Also a core reason why LLM-as-a-judge works so well: quality evaluation is vastly easier than generation.

<not this paper approach>

This is one of the key prompting in a lot of Enterprise cases. You can currently prompt LLMs to add a confidence score along with their responses.

Especially when you are using LLMs for downstream NLP tasks.

The confidence score can be a great indicator also for applying a two-tier model approach!

In my experience self validation performance and confidence ratings are both poor. I think the problem is that these sorts of formats just aren’t that common in the training data. What does help is to as a series of structured questions pertaining to quality and to aggregate those, but it’s still often not that helpful.
The challenge is the confidence scores can often be confabulations themselves.
What is a confidence score like that based on?
yeah but they say in the paper that isn't very accurate, they suggest a specifically fine-tuned model for it, not just prompting it for a score?
I admit I'm a bit confused by the reward function, as given it seems to provide the same score independent of correctness due to the squaring? And I think even if that's a mistake and it's supposed to be negative for incorrect answers, a policy that optimizes for that reward is to output 1 for anything with less than a 50% chance of being true and 10 for anything over 50%. Is that how RL is typically done?
it is nice that you posted datasets