about
Visual prompt engineering for video models (arxiv.org)
4 points by root-parent 66 days ago | hide | past | pdf | discuss on HN

In plain words: Before asking a video model a question, an image editor rewrites the picture—turning a sketch into a realistic scene—so the model sees an easier version. This improved reasoning across tasks, often beating text prompt tweaks or running the model longer before answering.

Abstract

In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential technique for improving language model performance. Since video models are currently becoming foundation models for visual tasks (e.g., visual reasoning), we here ask whether they similarly benefit from visual prompt engineering: automatically modifying the task image to improve model performance. For example, for a visual physics reasoning task ("Where does the ball land, after passing a set of obstacles?"), an abstract sketch-like scene can be turned into a photorealistic version with a simple call to an image editing model. We find that visual prompt engineering, or VIPE for short, improves video reasoning performance across tasks. In fact, for video models, visual prompt engineering can be even more effective than classic text-based prompt engineering or test-time scaling. Ultimately, just as text-based prompt engineering systematically improves language model performance, visual prompt engineering can serve as a simple, compute-efficient approach to elicit better visual reasoning performance from video models. Example videos on our project page at https://visual-prompt-engineering.github.io/.

Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer, Neha Kalibhat, Zi Wang, Mani Malek, Oyvind Tafjord, Kevin Swersky, Been Kim, Priyank Jaini
arXiv:2607.25537 · cs.CV, cs.AI · submitted Jul 28, 2026
abstract · pdf · html

add comment on HN