In plain words: By reading the hidden signals inside a language model, the study checks whether it knows where a fact sits in a long passage. It does know, yet often fails to say it — the location is stored but dropped when writing the answer.
Abstract
Large Language Models (LLMs) exhibit positional bias, struggling to utilize information from the middle or end of long contexts. Our study explores LLMs' long-context reasoning by probing their hidden representations. We find that while LLMs encode the position of target information, they often fail to leverage this in generating accurate responses. This reveals a disconnect between information retrieval and utilization, a "know but don't tell" phenomenon. We further analyze the relationship between extraction time and final accuracy, offering insights into the underlying mechanics of transformer models.
Taiming Lu, Muhan Gao, Kuai Yu, Adam Byerly, Daniel Khashabi
arXiv:2406.14673 · cs.CL · submitted Jun 20, 2024 · updated Oct 4, 2024
abstract · pdf · html