In plain words: A text generator is given a random hidden variable, learned without labels, that shapes each thing it writes. Compared with the usual generator that has no such variable, it does substantially better on later tasks.
Abstract
We propose an extension of the decoder Transformer that conditions its generative process on random latent variables which are learned without supervision thanks to a variational procedure. Experimental evaluations show that allowing such a conditioning translates into substantial improvements on downstream tasks.
François Fleuret
arXiv:2510.17558 · cs.LG · submitted Oct 20, 2025
abstract · pdf · html