In plain words: Fourteen deliberately tricky text prompts were given to an image generator to test its common sense, reasoning, and grasp of complex sentences, with ten pictures made for each. Five prompts produced at least one perfect picture, but no prompt got ten perfect ones.
Abstract · A very preliminary analysis of DALL-E 2
The DALL-E 2 system generates original synthetic images corresponding to an input text as caption. We report here on the outcome of fourteen tests of this system designed to assess its common sense, reasoning and ability to understand complex texts. All of our prompts were intentionally much more challenging than the typical ones that have been showcased in recent weeks. Nevertheless, for 5 out of the 14 prompts, at least one of the ten images fully satisfied our requests. On the other hand, on no prompt did all of the ten images satisfy our requests.
Gary Marcus, Ernest Davis, Scott Aaronson
arXiv:2204.13807 · cs.CV, cs.AI · submitted Apr 25, 2022 · updated May 2, 2022
abstract · pdf
> To the extent that the goal is to develop artificial intelligence that can be trusted in safety-critical applications (Marcus & Davis, 2019), a much higher standard must be applied.
I have seen no mention of that being the intended use of DALL-E (2). In fact, the most common use case I've seen described is in replacing Fiverr-type tasks: quick graphic design.
That said, the results are actually encouraging given that it _wasn't_ designed to succeed here:
> Nevertheless, for 5 out of the 14 prompts, at least one of the ten images fully satisfied our requests.
Some of the authors' interpretations could be argued against, as well. For instance, in example 10, "An old man is talking to his parents":
> In none of these images did DALL-E successfully infer that image should show an old man with two even older people
Several of the images appear to show exactly that? How is the author judging "even older"?