about
Comparing Humans, GPT-4, and GPT-4V on Abstraction and Reasoning Tasks (arxiv.org)
1 point by georgehill on Nov 17, 2023 | hide | past | pdf | discuss on HN

In plain words: They tested text-only and image-reading GPT-4 on puzzles requiring simple concepts like objects, counts, and shapes, giving the model one worked example first instead of none. Both versions still fell short of human-level abstraction, even with that extra example.

Abstract · Comparing Humans, GPT-4, and GPT-4V On Abstraction and Reasoning Tasks

We explore the abstract reasoning abilities of text-only and multimodal versions of GPT-4, using the ConceptARC benchmark [10], which is designed to evaluate robust understanding and reasoning with core-knowledge concepts. We extend the work of Moskvichev et al. [10] by evaluating GPT-4 on more detailed, one-shot prompting (rather than simple, zero-shot prompts) with text versions of ConceptARC tasks, and by evaluating GPT-4V, the multimodal version of GPT-4, on zero- and one-shot prompts using image versions of the simplest tasks. Our experimental results support the conclusion that neither version of GPT-4 has developed robust abstraction abilities at humanlike levels.

Melanie Mitchell, Alessandro B. Palmarini, Arseny Moskvichev
arXiv:2311.09247 · cs.AI, cs.LG · submitted Nov 14, 2023 · updated Dec 11, 2023
abstract · pdf · html · Corrected Figure 3 (extra spaces were replaced by commas, which were lost in original formatting)

add comment on HN
Also discussed: Nov 2023 (217 points, 177 comments)