about
NGPT: Normalized Transformer with Representation Learning on the Hypersphere (arxiv.org)
4 points by programd on Oct 18, 2024 | hide | past | pdf | discuss on HN

In plain words: Every vector in the network is kept the same length, so each layer nudges a word's representation along a sphere's surface toward its prediction. It matches the normal version's accuracy in 4 to 20 times fewer training steps, with bigger gains on longer sequences.

Abstract · nGPT: Normalized Transformer with Representation Learning on the Hypersphere

We propose a novel neural network architecture, the normalized Transformer (nGPT) with representation learning on the hypersphere. In nGPT, all vectors forming the embeddings, MLP, attention matrices and hidden states are unit norm normalized. The input stream of tokens travels on the surface of a hypersphere, with each layer contributing a displacement towards the target output predictions. These displacements are defined by the MLP and attention blocks, whose vector components also reside on the same hypersphere. Experiments show that nGPT learns much faster, reducing the number of training steps required to achieve the same accuracy by a factor of 4 to 20, depending on the sequence length.

Ilya Loshchilov, Cheng-Ping Hsieh, Simeng Sun, Boris Ginsburg
arXiv:2410.01131 · cs.LG, cs.AI · submitted Oct 1, 2024 · updated Apr 23, 2025
abstract · pdf · html

add comment on HN