about
Step-Video-T2V: The Practice, Challenges, and Future of Video Foundation Model (arxiv.org)
41 points by limoce on Feb 17, 2025 | hide | past | pdf | 5 comments on HN

In plain words: This video generator turns English or Chinese text into clips up to 204 frames by shrinking the video data and removing noise, then a final pass teaches it to prefer clean frames. On a new prompt test it beat free and paid video tools.

Abstract · Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

We present Step-Video-T2V, a state-of-the-art text-to-video pre-trained model with 30B parameters and the ability to generate videos up to 204 frames in length. A deep compression Variational Autoencoder, Video-VAE, is designed for video generation tasks, achieving 16x16 spatial and 8x temporal compression ratios, while maintaining exceptional video reconstruction quality. User prompts are encoded using two bilingual text encoders to handle both English and Chinese. A DiT with 3D full attention is trained using Flow Matching and is employed to denoise input noise into latent frames. A video-based DPO approach, Video-DPO, is applied to reduce artifacts and improve the visual quality of the generated videos. We also detail our training strategies and share key observations and insights. Step-Video-T2V's performance is evaluated on a novel video generation benchmark, Step-Video-T2V-Eval, demonstrating its state-of-the-art text-to-video quality when compared with both open-source and commercial engines. Additionally, we discuss the limitations of current diffusion-based model paradigm and outline future directions for video foundation models. We make both Step-Video-T2V and Step-Video-T2V-Eval available at https://github.com/stepfun-ai/Step-Video-T2V. The online version can be accessed from https://yuewen.cn/videos as well. Our goal is to accelerate the innovation of video foundation models and empower video content creators.

Guoqing Ma, Haoyang Huang, Kun Yan, Liangyu Chen, Nan Duan, Shengming Yin, Changyi Wan, Ranchen Ming, Xiaoniu Song, Xing Chen, Yu Zhou, Deshan Sun, et al.
arXiv:2502.10248 · cs.CV, cs.CL · submitted Feb 14, 2025 · updated Feb 24, 2025
abstract · pdf · html · 36 pages, 14 figures

add comment on HN

Nice, they include some model weights. The video examples seem to have some temporal flickering.

Repo with examples: https://github.com/stepfun-ai/Step-Video-T2V

off-topic: i've never seen a paper with this many authors.:-)
Check out the CERN or any big project ones. Hundreds and hundreds.
I found this one with over a thousand autors: https://arxiv.org/abs/2502.10291
DeepSeek is another big one, "100 additional authors not shown"

https://arxiv.org/abs/2412.19437