In plain words: Large language models learn by crunching huge piles of word strings, while human language runs on a built-in mental system that builds nested structures from very little input and rejects impossible languages. Because of that gap, the models cannot explain how language works in the mind.
Abstract
Large Language Models are useless for linguistics, as they are probabilistic models that require a vast amount of data to analyse externalized strings of words. In contrast, human language is underpinned by a mind-internal computational system that recursively generates hierarchical thought structures. The language system grows with minimal external input and can readily distinguish between real language and impossible languages.
Johan J. Bolhuis, Andrea Moro, Stephen Crain, Sandiway Fong
arXiv:2512.13441 · cs.CL, q-bio.NC · submitted Dec 15, 2025
abstract · pdf · html