In plain words: Tiny classifiers read a model's inner states to spot made-up spans in its answers, instead of a human judge. On one task they matched or beat a human expert, hit 95% of their best score by layer four, but fail on other tasks.
Abstract
We design probes trained on the internal representations of a transformer language model to predict its hallucinatory behavior on three grounded generation tasks. To train the probes, we annotate for span-level hallucination on both sampled (organic) and manually edited (synthetic) reference outputs. Our probes are narrowly trained and we find that they are sensitive to their training domain: they generalize poorly from one task to another or from synthetic to organic hallucinations. However, on in-domain data, they can reliably detect hallucinations at many transformer layers, achieving 95% of their peak performance as early as layer 4. Here, probing proves accurate for evaluating hallucination, outperforming several contemporary baselines and even surpassing an expert human annotator in response-level detection F1. Similarly, on span-level labeling, probes are on par or better than the expert annotator on two out of three generation tasks. Overall, we find that probing is a feasible and efficient alternative to language model hallucination evaluation when model states are available.
Sky CH-Wang, Benjamin Van Durme, Jason Eisner, Chris Kedzie
arXiv:2312.17249 · cs.CL, cs.AI, cs.LG · submitted Dec 28, 2023 · updated Jun 8, 2024
abstract · pdf · html · ACL 2024 (Findings) Camera-Ready
All good but I wonder how this is supposed to scale? As they also do acknowledge:
"Our work is motivated by the need for efficient LLM hallucination evaluation. Though computationally efficient, one of the key limitations of probing is in the need for labeled in-domain data for probe training."