about
Language Models Are Injective and Hence Invertible (arxiv.org)
3 points by QueensGambit 344 days ago | hide | past | pdf | 1 comment on HN

In plain words: A language model never turns two different texts into the same hidden numbers, so the original words can be rebuilt exactly from inside the model. Proofs and billions of tests found zero overlaps, and a new algorithm recovers the text in one pass.

Abstract · Language Models are Injective and Hence Invertible

Transformer components such as non-linear activations and normalization are inherently non-injective, suggesting that different inputs could map to the same output and prevent exact recovery of the input from a model's representations. In this paper, we challenge this view. First, we prove mathematically that transformer language models mapping discrete input sequences to their corresponding sequence of continuous representations are injective and therefore lossless, a property established at initialization and preserved during training. Second, we confirm this result empirically through billions of collision tests on six state-of-the-art language models, and observe no collisions. Third, we operationalize injectivity: we introduce SipIt, the first algorithm that provably and efficiently reconstructs the exact input text from hidden activations, establishing linear-time guarantees and demonstrating exact invertibility in practice. Overall, our work establishes injectivity as a fundamental and exploitable property of language models, with direct implications for transparency, interpretability, and safe deployment.

Giorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli, Yannis Panagakis, Emanuele Rodolà
arXiv:2510.15511 · cs.LG, cs.AI · submitted Oct 17, 2025 · updated Mar 13, 2026
abstract · pdf · html

add comment on HN
Also discussed: Oct 2025 (231 points, 147 comments) · Oct 2025 (1 point, 2 comments) · Oct 2025 (4 points, 1 comment) · Oct 2025 (1 point, 0 comments) · Oct 2025 (1 point, 0 comments)

Interesting but there is something weird about this analysis b/c they seem to assume a discrete input space & a continuous output space of uncountable cardinality whereas the input space is necessarily finite dimensional (in the continuous case) & of finite cardinality in the discrete case. It would be more interesting if the analysis actually used finiteness of the output space b/c real numbers are an idealization & in theory it is possible to injectively embed any discrete set into the real numbers simply b/c of the logic of cardinal arithmetic.