In plain words: Instead of writing left to right, this model tags each output slot with its position so it can pick the order on the fly and fill in any missing pieces. It cut the generation steps by ten times across language, path, and flight tasks.
Abstract · σ-GPTs: A New Approach to Autoregressive Models
Autoregressive models, such as the GPT family, use a fixed order, usually left-to-right, to generate sequences. However, this is not a necessity. In this paper, we challenge this assumption and show that by simply adding a positional encoding for the output, this order can be modulated on-the-fly per-sample which offers key advantageous properties. It allows for the sampling of and conditioning on arbitrary subsets of tokens, and it also allows sampling in one shot multiple tokens dynamically according to a rejection strategy, leading to a sub-linear number of model evaluations. We evaluate our method across various domains, including language modeling, path-solving, and aircraft vertical rate prediction, decreasing the number of steps required for generation by an order of magnitude.
Arnaud Pannatier, Evann Courdier, François Fleuret
arXiv:2404.09562 · cs.LG, cs.AI · submitted Apr 15, 2024 · updated Jul 1, 2024
abstract · pdf · html · 23 pages, 7 figures, accepted at ECML/PKDD 2024
The authors randomly permute (i.e., shuffle) input tokens in training and add two positional encodings to each token: one with the token's position and another with the position of the token to be predicted. Otherwise, the model is a standard autoregressive GPT. The consequences of this seemingly "simple" modification are significant:
* The authors can prompt the trained model with part of a sequence and then decode the missing tokens, all at once, in parallel, regardless of order -- i.e., the model can in-fill in parallel.
* The authors can compute conditional probability densities for every missing token in a sequence, again in parallel, i.e., densities for all missing tokens at once.
* The authors propose a rejection-sampling method for generating in-fill tokens, again in parallel. Their method seems to work well in practice.
I've added this to my reading list. Thank you for sharing it on HN.