about
Traveling words: a geometric interpretation of transformers (arxiv.org)
80 points by d4rkp4ttern on Oct 5, 2023 | hide | past | pdf | 4 comments on HN

In plain words: Layer normalization pins each word's meaning to a sphere's surface, and attention slides it around as text is processed. Probing a small GPT-2 found clear matching between what words ask for and what others offer early on, while deeper heads focused on specific subjects.

Abstract · Traveling Words: A Geometric Interpretation of Transformers

Transformers have significantly advanced the field of natural language processing, but comprehending their internal mechanisms remains a challenge. In this paper, we introduce a novel geometric perspective that elucidates the inner mechanisms of transformer operations. Our primary contribution is illustrating how layer normalization confines the latent features to a hyper-sphere, subsequently enabling attention to mold the semantic representation of words on this surface. This geometric viewpoint seamlessly connects established properties such as iterative refinement and contextual embeddings. We validate our insights by probing a pre-trained 124M parameter GPT-2 model. Our findings reveal clear query-key attention patterns in early layers and build upon prior observations regarding the subject-specific nature of attention heads at deeper layers. Harnessing these geometric insights, we present an intuitive understanding of transformers, depicting them as processes that model the trajectory of word particles along the hyper-sphere.

Raul Molina
arXiv:2309.07315 · cs.CL, cs.AI, cs.LG · submitted Sep 13, 2023 · updated Sep 19, 2023
abstract · pdf · html

add comment on HN

I like this definition. Simple, clean, effective.

Maybe the start of a new field tackling 'travelling wordsman' (smith?) problems.

To anyone knowledgeable: where does geometric deep learning fit in here? Is this paper just another geometric view, or is it an attempt at a formalization for transformer mathematics? I don't see prominent GDL authors in this paper's references (Bronstein, Cohen, Bruna, Veličković, ...).
I think it could be applicable to developing a novel geometric theory if the dynamics of the word trajectories were better understood.
Geometric learning is more about group- and invariant theory. This paper only goes as far as linear algebra, as far as I see.