about
Learning to Teach Large Language Models Logical Reasoning (arxiv.org)
2 points by rntn on Dec 27, 2023 | hide | past | pdf | 1 comment on HN

In plain words: They tested how well language models judge links between events, like which happened first or caused what, and found the models often contradict themselves. Teaching the models event-relation logic with a new practice set made their answers more consistent across tasks.

Abstract · Improving Large Language Models in Event Relation Logical Prediction

Event relations are crucial for narrative understanding and reasoning. Governed by nuanced logic, event relation extraction (ERE) is a challenging task that demands thorough semantic understanding and rigorous logical reasoning. In this paper, we conduct an in-depth investigation to systematically explore the capability of LLMs in understanding and applying event relation logic. More in detail, we first investigate the deficiencies of LLMs in logical reasoning across different tasks. Our study reveals that LLMs are not logically consistent reasoners, which results in their suboptimal performance on tasks that need rigorous reasoning. To address this, we explore three different approaches to endow LLMs with event relation logic, and thus enable them to generate more coherent answers across various scenarios. Based on our approach, we also contribute a synthesized dataset (LLM-ERL) involving high-order reasoning for evaluation and fine-tuning. Extensive quantitative and qualitative analyses on different tasks also validate the effectiveness of our approaches and provide insights for solving practical tasks with LLMs in future work. Codes are available at https://github.com/chenmeiqii/Teach-LLM-LR.

Meiqi Chen, Yubo Ma, Kaitao Song, Yixin Cao, Yan Zhang, Dongsheng Li
arXiv:2310.09158 · cs.AI · submitted Oct 13, 2023 · updated Aug 9, 2024
abstract · pdf · html · ACL 2024

add comment on HN

Our study demonstrates that LLMs are not good reasoners in solving tasks with rigorous reasoning and will produce counterfactual answers

Too much "A", not enough "I".

Is unreliable software really all that useful?