about
Distinguishing Ignorance from Error in LLM Hallucinations (arxiv.org)
2 points by micrum on Dec 4, 2024 | hide | past | pdf | discuss on HN

In plain words: They split hallucinations into two kinds: the model never learned the answer, or it knows the answer but still says something wrong. The second kind is common across models and questions, and telling the two apart fixes errors better than treating all hallucinations alike.

Abstract

Large language models (LLMs) are susceptible to hallucinations -- factually incorrect outputs -- leading to a large body of work on detecting and mitigating such cases. We argue that it is important to distinguish between two types of hallucinations: ones where the model does not hold the correct answer in its parameters, which we term HK-, and ones where the model answers incorrectly despite having the required knowledge, termed HK+. We first find that HK+ hallucinations are prevalent and occur across models and datasets. Then, we demonstrate that distinguishing between these two cases is beneficial for mitigating hallucinations. Importantly, we show that different models hallucinate on different examples, which motivates constructing model-specific hallucination datasets for training detectors. Overall, our findings draw attention to classifying types of hallucinations and provide means to handle them more effectively. The code is available at https://github.com/technion-cs-nlp/hallucination-mitigation .

Adi Simhi, Jonathan Herzig, Idan Szpektor, Yonatan Belinkov
arXiv:2410.22071 · cs.CL · submitted Oct 29, 2024 · updated Feb 18, 2025
abstract · pdf · html

add comment on HN
Also discussed: Nov 2024 (2 points, 0 comments)