In plain words: A test gives a model the same mistake, blamed on the user or itself, to see if it fixes it. Models fixed the user-blamed version but missed the same one they made, a 64.5% blind spot: they can catch errors but freeze on their own.
Abstract · Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
Although large language models (LLMs) have transformed AI, they still make errors and follow unproductive reasoning paths. Self-correction is vital for safety-critical applications, but studying it requires disentangling activation failure from knowledge deficiency: when a model fails to correct an error, is it because it cannot, or because it does not? We introduce Self-Correction Bench, a controlled evaluation framework that isolates this distinction by injecting the same error as either an external (user-attributed) or internal (model-attributed) error, keeping all other context identical. Testing 14 open-source non-reasoning models reveals a 64.5% Self-Correction Blind Spot: models correct external errors but fail on identical internal ones, proving the capability exists but is not activated. On models' own naturally generated errors, a measurable share of what a model fails to catch in its own output is caught when the identical error is presented externally. We trace the cause to post-training data composition: supervised fine-tuning datasets lack error-correction sequences, and fine-tuning with as few as 5,306 such traces already reduces the blind spot by 76.0%. Mechanistically, we identify a transferable conversational-role direction in representation space that causally gates self-correction. Appending "Wait" requires no training yet reduces the blind spot by 89.3%, and operates through a nearly independent pathway, indicating that correction activation is not reducible to this single mechanism.
Ken Tsui
arXiv:2507.02778 · cs.CL, cs.AI, cs.LG · submitted Jul 3, 2025 · updated Aug 2, 2026
abstract · pdf · html · Accepted to COLM 2026
Quite an interesting result. Per my chats with one of the online models, this magic token "wait" is generally applicable across all models.