about
Walking to the Car Wash: The Salience Bias of LLMs in Commonsense Reasoning (arxiv.org)
2 points by theanonymousone 60 days ago | hide | past | pdf | 1 comment on HN

In plain words: They built 1,145 problems where a flashy number hides a commonsense rule, to see if AI chatbots spot the trap. Models often fell for it, yet they knew the rule: without the misleading setup they stated it, and simple prompting fixed much of it.

Abstract · Would You Walk to the Car Wash? Salience Bias in LLM Commonsense Reasoning

Despite advances in complex reasoning, large language models (LLMs) can prioritize explicit input conditions over implicit task prerequisites. In everyday commonsense reasoning, this can lead to a failure we term Salience Bias: salient but task-irrelevant details (e.g., numerical values) draw models into computation while they overlook the physical or commonsense prerequisites of the task. A critical open question is whether this failure reflects a genuine gap in commonsense knowledge or merely its suppression under misleading task framing. To investigate this, we construct the SaliTrap Benchmark, comprising 1,145 items in four trap dimensions. Evaluating 12 LLMs, we find substantial vulnerability across the tested models, with higher numerical distractor counts associated with lower trap-avoidance rates and trap recognition not always leading to avoidance. Further probing of sycophantic-compliance cases shows that the relevant commonsense can often be elicited when the original task framing is removed, suggesting a gap between recognizing a constraint and applying it during task execution. Building on this diagnosis, we further show that lightweight, inference-time prompting alone substantially closes the gap without any retraining. Our findings highlight the importance of applying commonsense constraints during task execution, and we release SaliTrap as a testbed for studying this gap. The codes are available at https://github.com/Wuzheng02/SaliTrap

Zheng Wu, Chenhao Xue, Shijie Zheng, Yijie Lu, Cheng Yang, Zhuosheng Zhang
arXiv:2607.28478 · cs.CL · submitted Jul 30, 2026 · updated Sep 27, 2026
abstract · pdf · html

add comment on HN

Hmm.

I wonder how similar this is to anchoring bias in humans?

https://en.wikipedia.org/wiki/Anchoring_effect