In plain words: A trick that translates one network's signals into another's maps 100 representations along a low-dimensional scale from word meanings through grammar and meaning to predicting future words. It predicts how well each matches brain activity, and its main axis traces the brain's language hierarchy.
Abstract · Low-Dimensional Structure in the Space of Language Representations is Reflected in Brain Responses
How related are the representations learned by neural language models, translation models, and language tagging tasks? We answer this question by adapting an encoder-decoder transfer learning method from computer vision to investigate the structure among 100 different feature spaces extracted from hidden representations of various networks trained on language tasks. This method reveals a low-dimensional structure where language models and translation models smoothly interpolate between word embeddings, syntactic and semantic tasks, and future word embeddings. We call this low-dimensional structure a language representation embedding because it encodes the relationships between representations needed to process language for a variety of NLP tasks. We find that this representation embedding can predict how well each individual feature space maps to human brain responses to natural language stimuli recorded using fMRI. Additionally, we find that the principal dimension of this structure can be used to create a metric which highlights the brain's natural language processing hierarchy. This suggests that the embedding captures some part of the brain's natural language representation structure.
Richard Antonello, Javier Turek, Vy Vo, Alexander Huth
arXiv:2106.05426 · cs.CL, cs.LG · submitted Jun 9, 2021 · updated Dec 10, 2025
abstract · pdf · html · Accepted to the Advances in Neural Information Processing Systems 34 (2021) Revised to include voxel selection details
* There is low-dimensional structure within the space of representations learned by language models.
* This low-dimensional structure is reflected in brain responses predicted by model embeddings -- specifically, they say they can recover known language processing hierarchies from the principal dimension of representations, and that the embeddings can be used to predict which representations map well to each area in the brain.
I'm going to read the paper carefully.