about
Improved GUI Grounding via Iterative Narrowing (arxiv.org)
2 points by sandwichsphinx on Nov 22, 2024 | hide | past | pdf | discuss on HN

In plain words: To help AI assistants click the right button on a screen, this method repeatedly zooms in on the area the model picks, narrowing to the exact spot. It beat the usual one-shot approach for both general and screen-trained models across many interface types.

Abstract

Graphical User Interface (GUI) grounding plays a crucial role in enhancing the capabilities of Vision-Language Model (VLM) agents. While general VLMs, such as GPT-4V, demonstrate strong performance across various tasks, their proficiency in GUI grounding remains suboptimal. Recent studies have focused on fine-tuning these models specifically for zero-shot GUI grounding, yielding significant improvements over baseline performance. We introduce a visual prompting framework that employs an iterative narrowing mechanism to further improve the performance of both general and fine-tuned models in GUI grounding. For evaluation, we tested our method on a comprehensive benchmark comprising various UI platforms and provided the code to reproduce our results.

Anthony Nguyen
arXiv:2411.13591 · cs.CV, cs.AI, cs.CL · submitted Nov 18, 2024 · updated Sep 11, 2025
abstract · pdf · html · Code available at https://github.com/ant-8/GUI-Grounding-via-Iterative-Narrowing

add comment on HN