In plain words: A new text generator scrambles words two ways—hiding them or swapping in random ones—so it can rewrite earlier words, not just fill blanks. At equal training cost it matched the best diffusion language models and could fix its own mistakes.
Abstract
While state-of-the-art language models achieve impressive results through next-token prediction, they have inherent limitations such as the inability to revise already generated tokens. This has prompted exploration of alternative approaches such as discrete diffusion. However, masked diffusion, which has emerged as a popular choice due to its simplicity and effectiveness, reintroduces this inability to revise words. To overcome this, we generalize masked diffusion, deriving a new family of general interpolating discrete diffusion (GIDD) which offers greater flexibility in the design of the noising processes. Leveraging a novel diffusion ELBO, we achieve compute-matched state-of-the-art performance in diffusion language modeling. Exploiting GIDD's flexibility, we explore a hybrid approach combining masking and uniform noise, leading to improved sample quality and unlocking the ability for the model to correct its own mistakes, an area where autoregressive models notoriously have struggled. Code: https://github.com/dvruette/gidd/
Dimitri von Rütte, Janis Fluri, Yuhui Ding, Antonio Orvieto, Bernhard Schölkopf, Thomas Hofmann
arXiv:2503.04482 · cs.CL, cs.AI, cs.LG · submitted Mar 6, 2025 · updated Jun 9, 2025
abstract · pdf · html · Published at ICML 2025; Code available at https://github.com/dvruette/gidd