about
Exploring Hidden Reasoning Process of Large Language Models by Misleading Them (arxiv.org)
8 points by belter on Mar 23, 2025 | hide | past | pdf | discuss on HN

In plain words: Models were taught deliberately wrong math and logic rules, then tested on new problem types to see if they would follow them. They did, even on word problems and everyday reasoning, suggesting the model pulls out a rule first, then reasons with it.

Abstract · Exploring the Hidden Reasoning Process of Large Language Models by Misleading Them

Large language models (LLMs) have been able to perform various forms of reasoning tasks in a wide range of scenarios, but are they truly engaging in task abstraction and rule-based reasoning beyond mere memorization? To answer this question, we propose a novel experimental approach, Misleading Fine-Tuning (MisFT), to examine whether LLMs perform abstract reasoning by altering their original understanding of fundamental rules. In particular, by constructing datasets with math expressions or logical formulas that contradict correct principles, we fine-tune the model to learn those contradictory rules and assess its generalization ability on unseen test domains. Through a series of experiments, we find that current LLMs are capable of applying contradictory rules to solve practical math word problems and natural language reasoning tasks, implying the presence of an internal mechanism in LLMs that abstracts before reasoning.

Guanyu Chen, Peiyang Wang, Yizhou Jiang, Yuqian Liu, Chujie Zhao, Ying Fang, Tianren Zhang, Feng Chen
arXiv:2503.16401 · cs.LG · submitted Mar 20, 2025 · updated Nov 20, 2025
abstract · pdf · html

add comment on HN