about
V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs (arxiv.org)
58 points by jonbaer on Jan 16, 2024 | hide | past | pdf | 4 comments on HN

In plain words: The system lets the language model decide where to look in a big, crowded picture, zooming in on the right spots before answering. It finds small visual details more reliably than today's image models, which scan the whole picture at once.

Abstract

When we look around and perform complex tasks, how we see and selectively process what we see is crucial. However, the lack of this visual search mechanism in current multimodal LLMs (MLLMs) hinders their ability to focus on important visual details, especially when handling high-resolution and visually crowded images. To address this, we introduce V*, an LLM-guided visual search mechanism that employs the world knowledge in LLMs for efficient visual querying. When combined with an MLLM, this mechanism enhances collaborative reasoning, contextual understanding, and precise targeting of specific visual elements. This integration results in a new MLLM meta-architecture, named Show, sEArch, and TelL (SEAL). We further create V*Bench, a benchmark specifically designed to evaluate MLLMs in their ability to process high-resolution images and focus on visual details. Our study highlights the necessity of incorporating visual search capabilities into multimodal systems. The code is available https://github.com/penghao-wu/vstar.

Penghao Wu, Saining Xie
arXiv:2312.14135 · cs.CV · submitted Dec 21, 2023 · updated Dec 26, 2023
abstract · pdf · html · Project page with code: https://vstar-seal.github.io/

add comment on HN

I've actually been thinking recently of the embodiment approach here in the fact that foveated rendering in VR/AR and eye tracking for 3D laptops means there's going to be much more eye tracking data that could feed into ML systems around what humans pay attention to given a paired image.

I'd imagine that data, combined with things like this, could really streamline visual processing, particularly for video.

Outside of facebook... who would be collecting that data in a systematic manner?
Mostly Meta, but it's not like they don't have AI ambitions.

Apple will soon.

Media companies will eventually have it too as they roll out apps for the hardware.

I don't think it will take much data relative to overall visual training data. As a supplementary layer, a little bit goes a long way.

Interesting, I wonder if there's a similar approach that can be taken for complex tasks in the "multimodal" virtual environment of a computer desktop or phone, working across/between multiple applications and web pages with various inputs and interactions. Taking a more general approach that might be closer to how humans work then ACT-1 and the like.