In plain words: Allegro is an open video generator that turns text into smooth, consistent clips, released with the full recipe for data, training, and testing. In a user study it beat other open models and most commercial ones, ranking just behind Hailuo and Kling.
Abstract
Significant advancements have been made in the field of video generation, with the open-source community contributing a wealth of research papers and tools for training high-quality models. However, despite these efforts, the available information and resources remain insufficient for achieving commercial-level performance. In this report, we open the black box and introduce $\textbf{Allegro}$, an advanced video generation model that excels in both quality and temporal consistency. We also highlight the current limitations in the field and present a comprehensive methodology for training high-performance, commercial-level video generation models, addressing key aspects such as data, model architecture, training pipeline, and evaluation. Our user study shows that Allegro surpasses existing open-source models and most commercial models, ranking just behind Hailuo and Kling. Code: https://github.com/rhymes-ai/Allegro , Model: https://huggingface.co/rhymes-ai/Allegro , Gallery: https://rhymes.ai/allegro_gallery .
Yuan Zhou, Qiuyue Wang, Yuxuan Cai, Huan Yang
arXiv:2410.15458 · cs.CV · submitted Oct 20, 2024
abstract · pdf · html