In plain words: They measure how densely a model's attention steps link words, and show this density sets how many independent directions of information reach the layers that build meaning. Theory and toy tests find more directions means more reasoning power, matching why recent reasoning tricks help.
Abstract
The advancement of large language models (LLMs) for real-world applications hinges critically on enhancing their reasoning capabilities. In this work, we explore the reasoning abilities of large language models (LLMs) through their geometrical understanding. We establish a connection between the expressive power of LLMs and the density of their self-attention graphs. Our analysis demonstrates that the density of these graphs defines the intrinsic dimension of the inputs to the MLP blocks. We demonstrate through theoretical analysis and toy examples that a higher intrinsic dimension implies a greater expressive capacity of the LLM. We further provide empirical evidence linking this geometric framework to recent advancements in methods aimed at enhancing the reasoning capabilities of LLMs.
Romain Cosentino, Sarath Shekkizhar
arXiv:2407.02678 · cs.AI, cs.CL · submitted Jul 2, 2024
abstract · pdf · html
In the middle... AI doesn't work very well.
If an AI writes a multi-step plan, where the pieces have to fit together, I've found it goes off the rails. Parts 1 and 3 of a 4-part plan are fine. So is part 2. However they don't fit together! AI has no concept of "these four parts have to be closely connected, building a whole". It just builds from A to B in four steps... but taking two different paths and stitching the pieces together poorly.