In plain words: The study checks whether large language models can reason, plan, and fix their own wrong answers by reviewing them. It finds no evidence that this self-checking trick works for these models, unlike for people.
Abstract · Can Large Language Models Reason and Plan?
While humans sometimes do show the capability of correcting their own erroneous guesses with self-critiquing, there seems to be no basis for that assumption in the case of LLMs.
Subbarao Kambhampati
arXiv:2403.04121 · cs.AI, cs.CL, cs.LG · submitted Mar 7, 2024 · updated Mar 8, 2024
abstract · pdf · html · arXiv admin note: text overlap with arXiv:2402.01817 (v2 add creative commons attribution to Figure 2 graphic)