In plain words: They tested whether giving a language model two worked examples with step-by-step reasoning helps it figure out what people believe and want. Models trained with human feedback all improved, and GPT-4 jumped from nearly 80% correct without examples to perfect accuracy with them.
Abstract
Large language models (LLMs) excel in many tasks in 2023, but they still face challenges in complex reasoning. Theory-of-mind (ToM) tasks, which require understanding agents' beliefs, goals, and mental states, are essential for common-sense reasoning involving humans, making it crucial to enhance LLM performance in this area. This study measures the ToM performance of GPT-4 and three GPT-3.5 variants (Davinci-2, Davinci-3, GPT-3.5-Turbo), and investigates the effectiveness of in-context learning in improving their ToM comprehension. We evaluated prompts featuring two-shot chain of thought reasoning and step-by-step thinking instructions. We found that LLMs trained with Reinforcement Learning from Human Feedback (RLHF) (all models excluding Davinci-2) improved their ToM accuracy via in-context learning. GPT-4 performed best in zero-shot settings, reaching nearly 80% ToM accuracy, but still fell short of the 87% human accuracy on the test set. However, when supplied with prompts for in-context learning, all RLHF-trained LLMs exceeded 80% ToM accuracy, with GPT-4 reaching 100%. These results demonstrate that appropriate prompting enhances LLM ToM reasoning, and they underscore the context-dependent nature of LLM cognitive capacities.
Shima Rahimi Moghaddam, Christopher J. Honey
arXiv:2304.11490 · cs.AI, cs.CL · submitted Apr 22, 2023 · updated Apr 26, 2023
abstract · pdf · 27 pages, 4 main figures, 2 supplementary figures
GPT-4 exceeded average adult human theory of mind when the simple prompt "Let's think step by step:" was added. It got every question right when the prompt included a couple of example puzzles.
Average human ability was from the following source: "To measure humans’ performance in ToM and Photo scenarios, we recruited 125 online participants through the Qualtrics platform. Participants were 18 to 65 years old, native English speakers, and located in the United States."
It looks like we will need more new benchmarks by the time GPT-5 comes!