about
Position Information Emerges in Causal Transformers Without Positional Encoding (arxiv.org)
3 points by PaulHoule on Jan 25, 2025 | hide | past | pdf | discuss on HN

In plain words: Transformers that only let each word see words before it can still tell where words sit, even with no position tags. Nearby words end up with more similar number codes than distant ones, and this happens even before training.

Abstract · Position Information Emerges in Causal Transformers Without Positional Encodings via Similarity of Nearby Embeddings

Transformers with causal attention can solve tasks that require positional information without using positional encodings. In this work, we propose and investigate a new hypothesis about how positional information can be stored without using explicit positional encoding. We observe that nearby embeddings are more similar to each other than faraway embeddings, allowing the transformer to potentially reconstruct the positions of tokens. We show that this pattern can occur in both the trained and the randomly initialized Transformer models with causal attention and no positional encodings over a common range of hyperparameters.

Chunsheng Zuo, Pavel Guerzhoy, Michael Guerzhoy
arXiv:2501.00073 · cs.CL, cs.LG · submitted Dec 30, 2024
abstract · pdf · html · Forthcoming at the International Conference on Computational Linguistics 2025 (COLING 2025)

add comment on HN