In plain words: Layer normalization pins each word's meaning to a sphere's surface, and attention slides it around as text is processed. Probing a small GPT-2 found clear matching between what words ask for and what others offer early on, while deeper heads focused on specific subjects.
Abstract · Traveling Words: A Geometric Interpretation of Transformers
Transformers have significantly advanced the field of natural language processing, but comprehending their internal mechanisms remains a challenge. In this paper, we introduce a novel geometric perspective that elucidates the inner mechanisms of transformer operations. Our primary contribution is illustrating how layer normalization confines the latent features to a hyper-sphere, subsequently enabling attention to mold the semantic representation of words on this surface. This geometric viewpoint seamlessly connects established properties such as iterative refinement and contextual embeddings. We validate our insights by probing a pre-trained 124M parameter GPT-2 model. Our findings reveal clear query-key attention patterns in early layers and build upon prior observations regarding the subject-specific nature of attention heads at deeper layers. Harnessing these geometric insights, we present an intuitive understanding of transformers, depicting them as processes that model the trajectory of word particles along the hyper-sphere.
Raul Molina
arXiv:2309.07315 · cs.CL, cs.AI, cs.LG · submitted Sep 13, 2023 · updated Sep 19, 2023
abstract · pdf · html
Maybe the start of a new field tackling 'travelling wordsman' (smith?) problems.