about
Self-Correction Bench: Revealing and Addressing LLM Self-Correction Blind Spot (arxiv.org)
1 point by yubblegum 360 days ago | hide | past | pdf | 2 comments on HN

In plain words: A test gives a model the same mistake, blamed on the user or itself, to see if it fixes it. Models fixed the user-blamed version but missed the same one they made, a 64.5% blind spot: they can catch errors but freeze on their own.

Abstract · Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

Although large language models (LLMs) have transformed AI, they still make errors and follow unproductive reasoning paths. Self-correction is vital for safety-critical applications, but studying it requires disentangling activation failure from knowledge deficiency: when a model fails to correct an error, is it because it cannot, or because it does not? We introduce Self-Correction Bench, a controlled evaluation framework that isolates this distinction by injecting the same error as either an external (user-attributed) or internal (model-attributed) error, keeping all other context identical. Testing 14 open-source non-reasoning models reveals a 64.5% Self-Correction Blind Spot: models correct external errors but fail on identical internal ones, proving the capability exists but is not activated. On models' own naturally generated errors, a measurable share of what a model fails to catch in its own output is caught when the identical error is presented externally. We trace the cause to post-training data composition: supervised fine-tuning datasets lack error-correction sequences, and fine-tuning with as few as 5,306 such traces already reduces the blind spot by 76.0%. Mechanistically, we identify a transferable conversational-role direction in representation space that causally gates self-correction. Appending "Wait" requires no training yet reduces the blind spot by 89.3%, and operates through a nearly independent pathway, indicating that correction activation is not reducible to this single mechanism.

Ken Tsui
arXiv:2507.02778 · cs.CL, cs.AI, cs.LG · submitted Jul 3, 2025 · updated Aug 2, 2026
abstract · pdf · html · Accepted to COLM 2026

add comment on HN

Note: Actual title is Self-Correction Bench: Revealing and Addressing the Self-Correction Blind Spot in LLMs, Ken Tsui, 2025.

Quite an interesting result. Per my chats with one of the online models, this magic token "wait" is generally applicable across all models.

“We uncover a systematic failure: LLMs cannot correct errors in their own outputs while successfully correcting identical errors from external sources - a limitation we term the Self-Correction Blind Spot.”