about
Large Language Diffusion Models (arxiv.org)
7 points by kadushka on Feb 23, 2025 | hide | past | pdf | 3 comments on HN

In plain words: Instead of writing words left to right, this model learns to fill in masked-out words, gradually unmasking a whole passage at once. At 8 billion parameters it matched same-size left-to-right models on math, code, and general tasks, and beat GPT-4o at completing poems backwards.

Abstract

The capabilities of large language models (LLMs) are widely regarded as relying on autoregressive models (ARMs). We challenge this notion by introducing LLaDA, a diffusion model trained from scratch under the pre-training and supervised fine-tuning (SFT) paradigm. LLaDA employs a forward data masking process and a reverse generation process, parameterized by a Transformer to predict masked tokens. It provides a principled generative approach for probabilistic inference by optimizing a likelihood lower bound. Across extensive benchmarks on general tasks, math, code, and so on, LLaDA demonstrates strong scalability and performs comparably to our self-constructed ARM baselines. Remarkably, LLaDA 8B is competitive with strong LLMs like LLaMA3 8B in in-context learning and, after SFT, exhibits impressive instruction-following abilities in case studies such as multi-turn dialogue. Moreover, LLaDA addresses the reversal curse, surpassing GPT-4o in a reversal poem completion task. Our findings show the promise of diffusion models for language modeling at scale and challenge the common assumption that core LLM capabilities discussed above inherently depend on ARMs. Project page and codes: https://ml-gsai.github.io/LLaDA-demo/.

Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, Chongxuan Li
arXiv:2502.09992 · cs.CL, cs.LG · submitted Feb 14, 2025 · updated Oct 18, 2025
abstract · pdf · html

add comment on HN
Also discussed: Feb 2026 (2 points, 0 comments) · Feb 2025 (2 points, 0 comments) · Feb 2025 (2 points, 0 comments)

https://news.ycombinator.com/item?id=43080189 previously posted to little attention.

I really think there ought to be more discussion of this paper.

copying from my previous comment: A first-generation diffusion model is beating LLama 3 in some areas, a model with a huge amount of tuning and improvement work. And it's from China again!

A whole new "tree" of development has opened up. With so many possibilities - traditional scaling laws, out-loud chain of thought, in-model layer-repeating chain of thought, and now diffusion models - it seems unlikely to me that LLMs are going to hit a wall that the river of technological progress cannot flow around.

I wonder how well they'll work at translation. The paper indicates that they're rather good at poetry.

Interesting times.

I'm still reading the paper, but my main question is how slow is the model compared to LLM of the same size. It seems like to get the best accuracy they need to set number of time steps to the number of tokens to be generated. Does it make it comparable in speed to an LLM?
Update: finished the paper, and as I suspected, there's a serious downside in speed and memory consumption. LLaDA model has to process the entire output sequence on every time step - without anything like KV cache. Also, full quadratic attention happens on the entire output sequence on every time step, which makes it unfeasible for sequence length longer than a few thousand tokens.