In plain words: Instead of answering in one shot, it chains reasoning steps, with one trained model picking the next step and another drawing the conclusion, so the steps cause the answer. On logic and science questions it beat one-shot answers, and its reasoning can be checked.
Abstract
Although contemporary large language models (LMs) demonstrate impressive question-answering capabilities, their answers are typically the product of a single call to the model. This entails an unwelcome degree of opacity and compromises performance, especially on problems that are inherently multi-step. To address these limitations, we show how LMs can be made to perform faithful multi-step reasoning via a process whose causal structure mirrors the underlying logical structure of the problem. Our approach works by chaining together reasoning steps, where each step results from calls to two fine-tuned LMs, one for selection and one for inference, to produce a valid reasoning trace. Our method carries out a beam search through the space of reasoning traces to improve reasoning quality. We demonstrate the effectiveness of our model on multi-step logical deduction and scientific question-answering, showing that it outperforms baselines on final answer accuracy, and generates humanly interpretable reasoning traces whose validity can be checked by the user.
Antonia Creswell, Murray Shanahan
arXiv:2208.14271 · cs.AI, cs.CL · submitted Aug 30, 2022
abstract · pdf · html
And what are we to make of the reasoning step that “All round things are loud. The bear is sound. Therefore the bear is soft.” The bear is sound? What does that even mean?
Is that a typo that should read, “the bear is round”? If so, why isn’t the conclusion that the bear is loud rather than soft? From the context we do not know that all round things are soft, only that all soft things are round.