In plain words: Three GPT versions answered 61 classic reasoning puzzles where people succeed or fall for the same tricks, and their answers were compared with a theory of human thinking. Newer versions matched human thinking more often but also made more human-like mistakes, from 18% to 33%.
Abstract · Humans in Humans Out: On GPT Converging Toward Common Sense in both Success and Failure
Increase in computational scale and fine-tuning has seen a dramatic improvement in the quality of outputs of large language models (LLMs) like GPT. Given that both GPT-3 and GPT-4 were trained on large quantities of human-generated text, we might ask to what extent their outputs reflect patterns of human thinking, both for correct and incorrect cases. The Erotetic Theory of Reason (ETR) provides a symbolic generative model of both human success and failure in thinking, across propositional, quantified, and probabilistic reasoning, as well as decision-making. We presented GPT-3, GPT-3.5, and GPT-4 with 61 central inference and judgment problems from a recent book-length presentation of ETR, consisting of experimentally verified data-points on human judgment and extrapolated data-points predicted by ETR, with correct inference patterns as well as fallacies and framing effects (the ETR61 benchmark). ETR61 includes classics like Wason's card task, illusory inferences, the decoy effect, and opportunity-cost neglect, among others. GPT-3 showed evidence of ETR-predicted outputs for 59% of these examples, rising to 77% in GPT-3.5 and 75% in GPT-4. Remarkably, the production of human-like fallacious judgments increased from 18% in GPT-3 to 33% in GPT-3.5 and 34% in GPT-4. This suggests that larger and more advanced LLMs may develop a tendency toward more human-like mistakes, as relevant thought patterns are inherent in human-produced training data. According to ETR, the same fundamental patterns are involved both in successful and unsuccessful ordinary reasoning, so that the "bad" cases could paradoxically be learned from the "good" cases. We further present preliminary evidence that ETR-inspired prompt engineering could reduce instances of these mistakes.
Philipp Koralus, Vincent Wang-Maścianica
arXiv:2303.17276 · cs.AI, cs.CL, cs.HC, cs.LG · submitted Mar 30, 2023
abstract · pdf · html · 10 pages
It was refreshing to have a discussion where the other party is actually listening to my arguments. Too bad it doesn't retain the info from the discussion though...
I realized that I could not have such a discussion with the vast majority of people in my industry, even if they were fully open minded; primarily because most people would not have so much knowledge at their disposal as ChatGPT had. I really felt that it had ALL of the knowledge on the subject. I just had to point it to show it the contradictions in the knowledge which it possessed.
It feels like it is able to reason from first principles and synthesize information but it often chooses to present the consensus view by default. You have to really draw out its knowledge in order to make it override the consensus view. Kind of reminds me of System 1 (fast) thinking versus System 2 (slow) thinking from the book "Thinking fast and slow."