In plain words: Language skill splits in two: knowing grammar and patterns, and using language to understand and act in the world, which brain science shows come from different systems. Language models handle the rules well but stumble on real-world use, often needing training or outside help.
Abstract
Large Language Models (LLMs) have come closest among all models to date to mastering human language, yet opinions about their linguistic and cognitive capabilities remain split. Here, we evaluate LLMs using a distinction between formal linguistic competence -- knowledge of linguistic rules and patterns -- and functional linguistic competence -- understanding and using language in the world. We ground this distinction in human neuroscience, which has shown that formal and functional competence rely on different neural mechanisms. Although LLMs are surprisingly good at formal competence, their performance on functional competence tasks remains spotty and often requires specialized fine-tuning and/or coupling with external modules. We posit that models that use language in human-like ways would need to master both of these competence types, which, in turn, could require the emergence of mechanisms specialized for formal linguistic competence, distinct from functional competence.
Kyle Mahowald, Anna A. Ivanova, Idan A. Blank, Nancy Kanwisher, Joshua B. Tenenbaum, Evelina Fedorenko
arXiv:2301.06627 · cs.CL, cs.AI · submitted Jan 16, 2023 · updated Mar 23, 2024
abstract · pdf · html · The two lead authors contributed equally to this work; published in "Trends in Cognnitive Sciences", March 2024
- the language network, which delivers formal linguistic competence - the multiple demand network, which provides reasoning ability - the default network, which tracks narratives above the clause level - the theory of mind network, which infers the mental state of another entity
This leads to their argument that a modular structure would lead to enhanced ability for an LLM to be both formally and functionally competent. (While LLMs currently exhibit human-level formal linguistic competence, their functional competence--the ability to navigate the real world through language--has room for improvement.)
Transformer models, they note, have degree of emergent modularity through "allowing different attention heads to attend to different input features."
I was wondering, is it possible to characterize the degree of emergent modularity in current systems?