In plain words: They tested text-only and image-reading GPT-4 on puzzles requiring simple concepts like objects, counts, and shapes, giving the model one worked example first instead of none. Both versions still fell short of human-level abstraction, even with that extra example.
Abstract · Comparing Humans, GPT-4, and GPT-4V On Abstraction and Reasoning Tasks
We explore the abstract reasoning abilities of text-only and multimodal versions of GPT-4, using the ConceptARC benchmark [10], which is designed to evaluate robust understanding and reasoning with core-knowledge concepts. We extend the work of Moskvichev et al. [10] by evaluating GPT-4 on more detailed, one-shot prompting (rather than simple, zero-shot prompts) with text versions of ConceptARC tasks, and by evaluating GPT-4V, the multimodal version of GPT-4, on zero- and one-shot prompts using image versions of the simplest tasks. Our experimental results support the conclusion that neither version of GPT-4 has developed robust abstraction abilities at humanlike levels.
Melanie Mitchell, Alessandro B. Palmarini, Arseny Moskvichev
arXiv:2311.09247 · cs.AI, cs.LG · submitted Nov 14, 2023 · updated Dec 11, 2023
abstract · pdf · html · Corrected Figure 3 (extra spaces were replaced by commas, which were lost in original formatting)
In the first batch of participants collected via Amazon Mechanical Turk, each received 11 problems (this batch also only had two “minimal Problems,” as opposed to three such problems for everyone else). However, preliminary data examination showed that some participants did not fully follow the study instructions and had to be excluded (see Section 5.2). In response, we made the screening criteria more strict (requiring a Master Worker qualification, 99% of HITs approved with at least 2000 HIT history, as opposed to 95% approval requirement in the first batch). Participants in all but the first batch were paid $10 upon completing the experiment. Participants in the first batch were paid $5. In all batches, the median pay-per-hour exceeded the U.S. minimal wage.
(Arseny Moskvichev et al)
So in conclusion, this isn't a random sample of (adult) humans, and the paper doesn't give standard deviations.
It would've been more interesting of they had sampled an age range of humans which we would place GPT-4 on rather than just 'it's not as good' which is all this paper can say, really.