about
AR-Diffusion: Auto-Regressive Diffusion Model for Text Generation (arxiv.org)
7 points by doener on May 28, 2025 | hide | past | pdf | 3 comments on HN

In plain words: A text generator blends word-by-word writing with diffusion's cleanup steps: words on the left finish cleaning up first, so later words can build on them. It beat other diffusion text models, matching their quality 100 to 600 times faster.

Abstract

Diffusion models have gained significant attention in the realm of image generation due to their exceptional performance. Their success has been recently expanded to text generation via generating all tokens within a sequence concurrently. However, natural language exhibits a far more pronounced sequential dependency in comparison to images, and the majority of existing language models are trained with a left-to-right auto-regressive approach. To account for the inherent sequential characteristic of natural language, we introduce Auto-Regressive Diffusion (AR-Diffusion). AR-Diffusion ensures that the generation of tokens on the right depends on the generated ones on the left, a mechanism achieved through employing a dynamic number of denoising steps that vary based on token position. This results in tokens on the left undergoing fewer denoising steps than those on the right, thereby enabling them to generate earlier and subsequently influence the generation of tokens on the right. In a series of experiments on various text generation tasks, including text summarization, machine translation, and common sense generation, AR-Diffusion clearly demonstrated its superiority over existing diffusion language models and that it can be $100\times\sim600\times$ faster when achieving comparable results. Our code is available at https://github.com/microsoft/ProphetNet/tree/master/AR-diffusion.

Tong Wu, Zhihao Fan, Xiao Liu, Yeyun Gong, Yelong Shen, Jian Jiao, Hai-Tao Zheng, Juntao Li, Zhongyu Wei, Jian Guo, Nan Duan, Weizhu Chen
arXiv:2305.09515 · cs.CL · submitted May 16, 2023 · updated Dec 13, 2023
abstract · pdf · html · Accept By NIPS 2023

add comment on HN

I hope we figure out how to scale up diffusion models. Gemini is a cool tech demo, and Inception is almost smart enough to be interesting. There's a lot of tasks that you just need a certain intelligence floor for, and after that getting smarter doesn't really help. For those, speed could be a huge differentiator.
I think that diffusion models might be a better way to generate code than autoregressive left-to-right generation
Title should probably contain [2023]