about
When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection (arxiv.org)
2 points by badmonster 313 days ago | hide | past | pdf | 1 comment on HN

In plain words: A video-captioning model checks its own descriptions against written rules and rewrites them, then trains on its best self-scored versions so it no longer needs repeated checking. It beat leading captioners on standard tests with more detail and fewer invented facts, at the same speed.

Abstract · VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection

Existing Video Detailed Captioning (VDC) methods predominantly rely on costly human annotations or distillation from powerful proprietary models, creating a dependency on external supervision. In this paper, we propose VDC-Agent, an autonomous self-evolving framework that empowers a single Multimodal Large Language Model (MLLM) to generate and refine high-quality captions through principle-guided self-reflection. To overcome the inference latency inherent in iterative refinement, we further propose to internalize this reflective capability into the model. Specifically, we construct VDC-Agent-19K, a preference dataset derived from the agent's self-scored trajectories, and introduce a Curriculum Direct Preference Optimization (DPO) strategy. This strategy leverages the quality gap between generated candidates to progressively align the model from easy to hard samples. Extensive experiments demonstrate that VDC-Agent achieves state-of-the-art performance on VDC and DREAM-1K benchmarks, generating captions with superior detail and faithfulness. Crucially, our internalization strategy retains the inference efficiency of the base model while significantly enhancing its generalization capabilities, as validated by both quantitative metrics and human evaluation.

Qiang Wang, Xinyuan Gao, Yuhang He, Jizhou Han, Jiangyang Li, SongLin Dong, Zhiheng Ma, Yihong Gong
arXiv:2511.19436 · cs.CV, cs.AI, cs.LG, cs.MM · submitted Nov 24, 2025 · updated Aug 11, 2026
abstract · pdf · html · Accepted to ECCV 2026. Project Page: https://vdcagent.github.io

add comment on HN

video captioning without human labels or large teacher models, improving accuracy and efficiency by self-generating and refining captions automatically.