about
Vista: A Test-Time Self-Improving Video Generation Agent (arxiv.org)
2 points by SweetSoftPillow 344 days ago | hide | past | pdf | discuss on HN

In plain words: It plans the scenes, generates several, picks the best, then critics judge picture, sound, and story before an agent rewrites the prompt and tries again. It beat the best generators in up to 60% of matchups; people preferred its videos 66.4% of the time.

Abstract · VISTA: A Test-Time Self-Improving Video Generation Agent

Despite rapid advances in text-to-video synthesis, generated video quality remains critically dependent on precise user prompts. Existing test-time optimization methods, successful in other domains, struggle with the multi-faceted nature of video. In this work, we introduce VISTA (Video Iterative Self-improvemenT Agent), a novel multi-agent system that autonomously improves video generation through refining prompts in an iterative loop. VISTA first decomposes a user idea into a structured temporal plan. After generation, the best video is identified through a robust pairwise tournament. This winning video is then critiqued by a trio of specialized agents focusing on visual, audio, and contextual fidelity. Finally, a reasoning agent synthesizes this feedback to introspectively rewrite and enhance the prompt for the next generation cycle. Experiments on single- and multi-scene video generation scenarios show that while prior methods yield inconsistent gains, VISTA consistently improves video quality and alignment with user intent, achieving up to 60% pairwise win rate against state-of-the-art baselines. Human evaluators concur, preferring VISTA outputs in 66.4% of comparisons.

Do Xuan Long, Xingchen Wan, Hootan Nakhost, Chen-Yu Lee, Tomas Pfister, Sercan Ö. Arık
arXiv:2510.15831 · cs.CV · submitted Oct 17, 2025
abstract · pdf · html

add comment on HN