In plain words: They tested whether four language models' internal sentence patterns record who did what to whom, by comparing how the models treated pairs of sentences. Unlike humans, the patterns tracked structure but not swapped roles, though a few attention units did capture them, more weakly.
Abstract
Large Language Models (LLMs) are commonly criticized for not understanding language. However, many critiques focus on cognitive abilities that, in humans, are distinct from language processing. Here, we instead study a kind of understanding tightly linked to language: inferring who did what to whom (thematic roles) in a sentence. Does the central training objective of LLMs-word prediction-result in sentence representations that capture thematic roles? In two experiments, we characterized sentence representations in four LLMs. In contrast to human similarity judgments, in LLMs the overall representational similarity of sentence pairs reflected syntactic similarity but not whether their agent and patient assignments were identical vs. reversed. Furthermore, we found little evidence that thematic role information was available in any subset of hidden units. However, some attention heads robustly captured thematic roles, independently of syntax. Therefore, LLMs can extract thematic roles but, relative to humans, this information influences their representations more weakly.
Joseph M. Denning, Xiaohan Hannah Guo, Bryor Snefjella, Idan A. Blank
arXiv:2504.16884 · cs.CL · submitted Apr 23, 2025 · updated Apr 25, 2025
abstract · pdf
Most likely because there is no mechanism in this thing that would allow for building spatial or relationship model between entities.