In plain words: They tested whether giving a language model two worked examples with step-by-step reasoning helps it figure out what people believe and want. Models trained with human feedback all improved, and GPT-4 jumped from nearly 80% correct without examples to perfect accuracy with them.
Abstract · Boosting Theory-of-Mind Performance in Large Language Models via Prompting
Large language models (LLMs) excel in many tasks in 2023, but they still face challenges in complex reasoning. Theory-of-mind (ToM) tasks, which require understanding agents' beliefs, goals, and mental states, are essential for common-sense reasoning involving humans, making it crucial to enhance LLM performance in this area. This study measures the ToM performance of GPT-4 and three GPT-3.5 variants (Davinci-2, Davinci-3, GPT-3.5-Turbo), and investigates the effectiveness of in-context learning in improving their ToM comprehension. We evaluated prompts featuring two-shot chain of thought reasoning and step-by-step thinking instructions. We found that LLMs trained with Reinforcement Learning from Human Feedback (RLHF) (all models excluding Davinci-2) improved their ToM accuracy via in-context learning. GPT-4 performed best in zero-shot settings, reaching nearly 80% ToM accuracy, but still fell short of the 87% human accuracy on the test set. However, when supplied with prompts for in-context learning, all RLHF-trained LLMs exceeded 80% ToM accuracy, with GPT-4 reaching 100%. These results demonstrate that appropriate prompting enhances LLM ToM reasoning, and they underscore the context-dependent nature of LLM cognitive capacities.
Shima Rahimi Moghaddam, Christopher J. Honey
arXiv:2304.11490 · cs.AI, cs.CL · submitted Apr 22, 2023 · updated Apr 26, 2023
abstract · pdf · 27 pages, 4 main figures, 2 supplementary figures
Some points addressed in the paper:
- Is the model emulating the reasoning of the 2-shot examples?
> The Davinci-3 and GPT-4 models experienced increases in ToM performance from all of the classes of CoT examples that we tested: Photo examples, Non-ToM Inferential examples, and ToM examples. The mean accuracy increases for each model and each type of CoT example are shown in Figure 4, while the accuracy changes for individual ToM questions are shown in Figure S.1. Prompting with Inferential and Photo examples boosted the models’ performance on ToM scenarios even though these in-context examples did not follow the same reasoning pattern as the ToM scenarios. Therefore, our analysis suggests that the benefit of prompting for boosting ToM performance is not due to merely overfitting to the specific set of reasoning steps shown in the CoT examples. Instead, the CoT examples appear to invoke a mode of output that involves step-by-step reasoning, which improves the accuracy across a range of tasks.
- Is the test data included in the training?
> The LLMs may have seen some ToM or Photo scenarios during their training phase, but data leakage is unlikely to affect our findings. First, our findings concern the change in performance arising from prompting, and the specific prompts used to obtain this performance change were novel materials generated for this study. Second, if the model performance relied solely on prior exposure to the training data, there should be little difference between zero-shot Photo and ToM performance (Figure 2), as these materials were published in the same documents; however, the zero-shot performance patterns were very different across Photo and ToM scenarios. Third, the LLM performance improvements arose when the models elaborated their reasoning step-by-step, and this elaborated reasoning was not part of the training data. Therefore, although some data leakage is possible, it is unlikely to affect our conclusions concerning the benefits of prompting.
Other highlights:
- With the prompting techniques 3.5-turbo got human level performance.
- In zero-shot GPT4 is already near human performance (90% of the human score)
- In zero-shot regime 3.5-turbo is worst than davinci-3, but much much better than it with prompting. This happens because turbo is too much cautious by default and often refuses to draw conclusions.