In plain words: A test pairs two images with two captions built from the exact same words in a different order, so you must match each picture to the right wording. Every vision-and-language model tested scored about as well as random guessing.
Abstract · Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality
We present a novel task and dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning, which we call Winoground. Given two images and two captions, the goal is to match them correctly - but crucially, both captions contain a completely identical set of words, only in a different order. The dataset was carefully hand-curated by expert annotators and is labeled with a rich set of fine-grained tags to assist in analyzing model performance. We probe a diverse range of state-of-the-art vision and language models and find that, surprisingly, none of them do much better than chance. Evidently, these models are not as skilled at visio-linguistic compositional reasoning as we might have hoped. We perform an extensive analysis to obtain insights into how future work might try to mitigate these models' shortcomings. We aim for Winoground to serve as a useful evaluation set for advancing the state of the art and driving further progress in the field. The dataset is available at https://huggingface.co/datasets/facebook/winoground.
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, Candace Ross
arXiv:2204.03162 · cs.CV, cs.CL · submitted Apr 7, 2022 · updated Apr 22, 2022
abstract · pdf · html · CVPR 2022
It's a real pity that the training data is so expensive to create, otherwise Google could feed all of YouTube into an AI and get it to understand physical reality as well as GPT-3 understands textual reality.
I guess some people would say that GPT-3 doesn't "understand" anything, it merely predicts tokens, but I think that GPT-3 does at least learn "sensible" patterns, so it would be nice if an AI could for example predict how humans (and cameras) move in a scene.
Perhaps being able to translate between textual and visual domains would help speed up machine training/learning, since humans often combine the two domains when learning new concepts too, but my guess is that AIs need to be able to symbolically manipulate abstract ideas to really bridge the gap between prediction and understanding.