about
Paris 2.0: Video diffusion model trained on decentralized, heterogeneous GPUs (arxiv.org)
7 points by royychacker 128 days ago | hide | past | pdf | 1 comment on HN

In plain words: A text-to-video generator trained by spreading the work across many separate computers instead of one giant GPU cluster. On the same data and compute it beat the single-cluster version, halving its video quality error score from 561 to 279.

Abstract · Paris 2.0: A Decentralized Diffusion Model for Video Generation

We present Paris 2.0, the first video generation model pre-trained through decentralized computation. Its training recipe builds upon Paris 1.0 (arXiv:2510.03434), the first ever open-weight Decentralized Diffusion Model (DDM), which showed that image generation can be trained without a monolithic GPU cluster. However, temporally coherent video generation had remained an open problem under decentralized training, and Paris 2.0 closes it. In low-resolution text-to-video training, against a monolithic model trained on the same data under a matched total compute budget, Paris 2.0 cuts Frechet Video Distance (FVD) from 561.04 to 279.01, a ~2.0x improvement, and lifts CLIP text-video similarity and aesthetic score.

Ali Rouzbayani, Bidhan Roy, Marcos Villagra, Zhiying Jiang
arXiv:2605.26064 · cs.CV, cs.LG · submitted May 25, 2026 · updated May 28, 2026
abstract · pdf · html · 6 pages, 5 figures

add comment on HN

Paris 2.0 is the first video generation model pre-trained through decentralized computation, building on Paris 1.0 (the first open-weight Decentralized Diffusion Model). It demonstrates that temporally coherent video generation is achievable without a monolithic GPU cluster. In low-resolution text-to-video training, it cuts Fréchet Video Distance (FVD) from 561 to 279 (~2x improvement) compared to a monolithic model trained on matched compute.