about
Explaining Why Hallucinate Large Language Models (arxiv.org)
1 point by gagan30 348 days ago | hide | past | pdf | 1 comment on HN

In plain words: By reading which concepts a model's internal states point to at each layer, and tracing which push which, this tool shows how an answer drifts toward a plausible but unsupported idea. Its explanations beat standard attribution and probing checks, and its score predicts failures.

Abstract · Distributional Semantics Tracing: A Framework for Explaining Hallucinations in Large Language Models

Hallucinations in large language models (LLMs) produce fluent continuations that are not supported by the prompt, especially under minimal contextual cues and ambiguity. We introduce Distributional Semantics Tracing (DST), a model-native method that builds layer-wise semantic maps at the answer position by decoding residual-stream states through the unembedding, selecting a compact top-$K$ concept set, and estimating directed concept-to-concept support via lightweight causal tracing. Using these traces, we test a representation-level hypothesis: hallucinations arise from correlation-driven representational drift across depth, where the residual stream is pulled toward a locally coherent but context-inconsistent concept neighborhood reinforced by training co-occurrences. On Racing Thoughts dataset, DST yields more faithful explanations than attribution, probing, and intervention baselines under an LLM-judge protocol, and the resulting Contextual Alignment Score (CAS) strongly predicts failures, supporting this drift hypothesis.

Gagan Bhatia, Somayajulu G Sripada, Kevin Allan, Jacobo Azcona
arXiv:2510.06107 · cs.CL, cs.AI, cs.CE · submitted Oct 7, 2025 · updated Mar 15, 2026
abstract · pdf · html

add comment on HN

Distributional Semantics Tracing: A Framework for Explaining Hallucinations in Large Language Models Gagan Bhatia, Somayajulu G Sripada, Kevin Allan, Jacobo Azcona Large Language Models (LLMs) are prone to hallucination, the generation of plausible yet factually incorrect statements. This work investigates the intrinsic, architectural origins of this failure mode through three primary contributions. First, to enable the reliable tracing of internal semantic failures, we propose Distributional Semantics Tracing (DST), a unified framework that integrates established interpretability techniques to produce a causal map of a model's reasoning, treating meaning as a function of context (distributional semantics). Second, we pinpoint the model's layer at which a hallucination becomes inevitable, identifying a specific commitment layer where a model's internal representations irreversibly diverge from factuality. Third, we identify the underlying mechanism for these failures. We observe a conflict between distinct computational pathways, which we interpret using the lens of dual-process theory: a fast, heuristic associative pathway (akin to System 1) and a slow, deliberate, contextual pathway (akin to System 2), leading to predictable failure modes such as Reasoning Shortcut Hijacks. Our framework's ability to quantify the coherence of the contextual pathway reveals a strong negative correlation () with hallucination rates, implying that these failures are predictable consequences of internal semantic weakness. The result is a mechanistic account of how, when, and why hallucinations occur within the Transformer architecture.