about
Dissociating language and thought in large language models (arxiv.org)
42 points by rntn on Sep 21, 2024 | hide | past | pdf | 4 comments on HN

In plain words: Language skill splits in two: knowing grammar and patterns, and using language to understand and act in the world, which brain science shows come from different systems. Language models handle the rules well but stumble on real-world use, often needing training or outside help.

Abstract

Large Language Models (LLMs) have come closest among all models to date to mastering human language, yet opinions about their linguistic and cognitive capabilities remain split. Here, we evaluate LLMs using a distinction between formal linguistic competence -- knowledge of linguistic rules and patterns -- and functional linguistic competence -- understanding and using language in the world. We ground this distinction in human neuroscience, which has shown that formal and functional competence rely on different neural mechanisms. Although LLMs are surprisingly good at formal competence, their performance on functional competence tasks remains spotty and often requires specialized fine-tuning and/or coupling with external modules. We posit that models that use language in human-like ways would need to master both of these competence types, which, in turn, could require the emergence of mechanisms specialized for formal linguistic competence, distinct from functional competence.

Kyle Mahowald, Anna A. Ivanova, Idan A. Blank, Nancy Kanwisher, Joshua B. Tenenbaum, Evelina Fedorenko
arXiv:2301.06627 · cs.CL, cs.AI · submitted Jan 16, 2023 · updated Mar 23, 2024
abstract · pdf · html · The two lead authors contributed equally to this work; published in "Trends in Cognnitive Sciences", March 2024

add comment on HN
Also discussed: Feb 2023 (2 points, 0 comments) · Jan 2023 (3 points, 0 comments) · Jan 2023 (2 points, 0 comments) · Jan 2023 (1 point, 1 comment)

The human brain, the authors argue, in fact uses multiple networks when interpreting and producing language. These include:

- the language network, which delivers formal linguistic competence - the multiple demand network, which provides reasoning ability - the default network, which tracks narratives above the clause level - the theory of mind network, which infers the mental state of another entity

This leads to their argument that a modular structure would lead to enhanced ability for an LLM to be both formally and functionally competent. (While LLMs currently exhibit human-level formal linguistic competence, their functional competence--the ability to navigate the real world through language--has room for improvement.)

Transformer models, they note, have degree of emergent modularity through "allowing different attention heads to attend to different input features."

I was wondering, is it possible to characterize the degree of emergent modularity in current systems?

One of the big limitations in LLMs is that they only have a single context window. People throw things that probably shouldn't be mixed together into the same context and hope for the best (e.g. system prompt, RAG context, user input, LLM output).

This is basically no different from a Turing machine going from one tape to multiple tapes. While in theory it doesn't make the Turing machine more powerful, it saves a whole lot of book keeping operations that are necessary to work around the limitations of a single tape.

Another limitation is the inability to seek to positions by moving the head back and forth to rewrite old data in the context.

I’m not sure if this is exactly what you are referring to, but Anthropic has done a lot of interpretability work on Claude, which they’ve published along with the famous "Golden Gate Claude".^1

"We also find more abstract features—responding to things like bugs in computer code, discussions of gender bias in professions, and conversations about keeping secrets."

1: https://www.anthropic.com/research/mapping-mind-language-mod...

How much of this modularity is it worth to insert in the architecture itself and how much should we let emerge from the training process itself?