In plain words: A language model is trained with rewards for correct outcomes so it writes out the reasoning behind people's risky choices in plain words. Unlike neural networks trained on behavior data that only guess the choice, it explains the thinking while still predicting decisions accurately.
Abstract · Using Reinforcement Learning to Train Large Language Models to Explain Human Decisions
A central goal of cognitive modeling is to develop models that not only predict human behavior but also provide insight into the underlying cognitive mechanisms. While neural network models trained on large-scale behavioral data often achieve strong predictive performance, they typically fall short in offering interpretable explanations of the cognitive processes they capture. In this work, we explore the potential of pretrained large language models (LLMs) to serve as dual-purpose cognitive models--capable of both accurate prediction and interpretable explanation in natural language. Specifically, we employ reinforcement learning with outcome-based rewards to guide LLMs toward generating explicit reasoning traces for explaining human risky choices. Our findings demonstrate that this approach produces high-quality explanations alongside strong quantitative predictions of human decisions.
Jian-Qiao Zhu, Hanbo Xie, Dilip Arumugam, Robert C. Wilson, Thomas L. Griffiths
arXiv:2505.11614 · cs.AI, cs.CL · submitted May 16, 2025 · updated Feb 1, 2026
abstract · pdf · html