In plain words: A model trained on mixed web text and images learns to read documents, describe pictures, answer questions, and solve puzzles from examples in its prompt. It does well on language, vision, and image-text tasks with no fine-tuning, and skills in one area help the others.
Abstract
A big convergence of language, multimodal perception, action, and world modeling is a key step toward artificial general intelligence. In this work, we introduce Kosmos-1, a Multimodal Large Language Model (MLLM) that can perceive general modalities, learn in context (i.e., few-shot), and follow instructions (i.e., zero-shot). Specifically, we train Kosmos-1 from scratch on web-scale multimodal corpora, including arbitrarily interleaved text and images, image-caption pairs, and text data. We evaluate various settings, including zero-shot, few-shot, and multimodal chain-of-thought prompting, on a wide range of tasks without any gradient updates or finetuning. Experimental results show that Kosmos-1 achieves impressive performance on (i) language understanding, generation, and even OCR-free NLP (directly fed with document images), (ii) perception-language tasks, including multimodal dialogue, image captioning, visual question answering, and (iii) vision tasks, such as image recognition with descriptions (specifying classification via text instructions). We also show that MLLMs can benefit from cross-modal transfer, i.e., transfer knowledge from language to multimodal, and from multimodal to language. In addition, we introduce a dataset of Raven IQ test, which diagnoses the nonverbal reasoning capability of MLLMs.
Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Barun Patra, Qiang Liu, Kriti Aggarwal, et al.
arXiv:2302.14045 · cs.CL, cs.CV · submitted Feb 27, 2023 · updated Mar 1, 2023
abstract · pdf · html
What's interesting is that it seems to actually lose information, as asking it to identify the studio that made WALL-E is beyond its capabilities, while asking it to describe the image (i.e. regenerating more closely something that was fed into it) and then processing on that text, is successful.
The "chain-of-thought" trick in LLMs I suspect underestimates the extent to which the interviewer is carrying water for the LLM's "reasoning" ability. The interviewer has a sense of what answer they want and will ask questions that produce further results that more easily prime the model to produce it. Reasoning supposes that these steps are carried out internally, but we see claims being made of reasoning when there is an external intelligence essentially directing the generation and combination of facts.
Another curious aspect is the flattening of 2D IQ test questions into linear format, which of course misses the point of the question in being able to reason spatially instead of linearly.