about
DiffusionBlocks: Block-Wise Neural Network Training (arxiv.org)
3 points by E-Reverance 257 days ago | hide | past | pdf | 5 comments on HN

In plain words: Residual connections are treated as steps in a denoising process, so each transformer block can be trained on its own instead of backpropagating through the whole network, cutting memory in proportion to the number of blocks. It matched end-to-end training on vision and generative tasks.

Abstract · DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation

End-to-end backpropagation requires storing activations throughout all layers, creating memory bottlenecks that limit model scalability. Existing block-wise training methods offer means to alleviate this problem, but they rely on ad-hoc local objectives and remain largely unexplored beyond classification tasks. We propose $\textit{DiffusionBlocks}$, a principled framework for transforming transformer-based networks into genuinely independent trainable blocks that maintain competitive performance with end-to-end training. Our key insight leverages the fact that residual connections naturally correspond to updates in a dynamical system. With minimal modifications to this system, we can convert the updates to those of a denoising process, where each block can be learned independently by leveraging the score matching objective. This independence enables training with gradients for only one block at a time, thereby reducing memory requirements in proportion to the number of blocks. Our experiments on a range of transformer architectures (vision, diffusion, autoregressive, recurrent-depth, and masked diffusion) demonstrate that DiffusionBlocks training matches the performance of end-to-end training while enabling scalable block-wise training on practical tasks beyond small-scale classification. DiffusionBlocks provides a theoretically grounded approach that successfully scales to modern generative tasks across diverse architectures. Code is available at https://github.com/SakanaAI/DiffusionBlocks .

Makoto Shing, Masanori Koyama, Takuya Akiba
arXiv:2506.14202 · cs.LG, cs.AI, stat.ML · submitted Jun 17, 2025 · updated Jun 12, 2026
abstract · pdf · html · To appear at the 14th International Conference on Learning Representations (ICLR 2026). v4: Fixed typos in experimental details (Appendix E.4)

add comment on HN

The same Sakana that couldn't validate experiments

https://techcrunch.com/2025/02/21/sakana-walks-back-claims-t...

Other teams did a better job and provided code

https://github.com/kuleshov-group/bd3lms

This is unrelated. They both use the word "block", but what they are referring to differs
"Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models"
Yes and? The paper I linked is about network weights, not the type of generative model