about
Superficial Beliefs in LLM Decision-Making (arxiv.org)
3 points by MediaSquirrel 115 days ago | hide | past | pdf | discuss on HN

In plain words: Models chose between profiles with graded traits, and a fit to earlier choices revealed which trait actually drove each pick. That trait predicted later choices well, but the models' stated reasons matched it only partly, so decisions are structured while explanations are shallow.

Abstract

We ask whether large language models (LLMs) merely imitate rationales when choosing between two options, or whether their choices reflect a systematic underlying decision structure. Using synthetic binary decision settings in which models choose between profiles defined by graded attributes, we compare the attribute a model says mattered most with the attribute that best explains its choice under a behavioural model fit to prior decisions. The behavioural model predicts held-out choices well, showing that model behaviour is systematically related to the visible attributes rather than being random. However, direct self-reports and a separate score-based judge recover the behaviourally inferred driver only partially. The resulting picture is neither one of arbitrary behaviour nor one of fully articulated belief - outputs are structured enough to support prediction, but explicit reasons track the recovered driver only imperfectly. This qualitative pattern persists across prompt-order and sampling perturbations, alternative behavioural models, targeted occlusion analyses, and structurally varied decision settings. We interpret this as evidence for ``superficial belief'' in LLM decision-making: models behave as if guided by probabilistic local priorities over attributes, while having only limited verbal access to the attributes that drive their decisions.

Gabriel Freedman, Francesca Toni
arXiv:2606.11016 · cs.AI · submitted Jun 9, 2026 · updated Sep 15, 2026
abstract · pdf · html · Published as a conference paper at COLM 2026

add comment on HN