about
Visual General Intelligence: A White Paper (arxiv.org)
2 points by Anon84 26 days ago | hide | past | pdf | discuss on HN

In plain words: Vision researchers lay out how learning from images, video, and 3D shapes could lead to artificial general intelligence — the kind that handles any task. It maps the principles, inputs, tests, and learning setups to pursue, without settling on one definition of visual intelligence.

Abstract

This paper reconsiders intelligence from a vision-centered perspective and examines whether intelligence emerging from visual experience and learning may provide a pathway toward AGI. In the language domain, beginning with the introduction of the Transformer architecture, the GPT series has demonstrated transfer to unseen tasks through autoregressive language modeling on web-scale text combined with aggressive scaling. This raises a natural question, namely, what capabilities and forms of intelligence can emerge from visual modalities such as images, videos, and geometry? In this paper, we discuss whether visual intelligence can serve as a pathway toward AGI, referred to in this paper as visual general intelligence (VGI), by bringing together contributors from diverse standpoints and affiliations. Our aim is not to offer a single definition of visual intelligence, but to clarify the principles that computer vision should pursue in the AGI era, the visual input modalities, the benchmarks, the learning paradigms, and the relationship between vision, when taken as the core, and other modalities such as language.

Hirokatsu Kataoka, Yoshihiro Fukuhara, Yonglong Tian, Shangzhe Wu, Oishi Deb, Ryousuke Yamada, Christian Rupprecht, Jianyuan Wang, Kohsuke Ide, Koichi Namekata, Xianzheng Ma, Yiming Chen, et al.
arXiv:2608.25924 · cs.CV · submitted Aug 26, 2026
abstract · pdf · html

add comment on HN