about
The Illusion of the Illusion of Thinking (arxiv.org)
12 points by jedisct1 on Jun 16, 2025 | hide | past | pdf | 1 comment on HN

In plain words: A recheck of puzzle tests that showed AI reasoning breaking down found the tests were flawed: some puzzles were impossible, others needed more output space than allowed. Asked for a formula instead of moves, models solved Tower of Hanoi cases previously scored as failures.

Abstract · Comment on The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity

Shojaee et al. (2025) report that Large Reasoning Models (LRMs) exhibit "accuracy collapse" on planning puzzles beyond certain complexity thresholds. We demonstrate that their findings primarily reflect experimental design limitations rather than fundamental reasoning failures. Our analysis reveals three critical issues: (1) Tower of Hanoi experiments risk exceeding model output token limits, with models explicitly acknowledging these constraints in their outputs; (2) The authors' automated evaluation framework fails to distinguish between reasoning failures and practical constraints, leading to misclassification of model capabilities; (3) Most concerningly, their River Crossing benchmarks include mathematically impossible instances for N > 5 due to insufficient boat capacity, yet models are scored as failures for not solving these unsolvable problems. When we control for these experimental artifacts, by requesting generating functions instead of exhaustive move lists, preliminary experiments across multiple models indicate high accuracy on Tower of Hanoi instances previously reported as complete failures. These findings highlight the importance of careful experimental design when evaluating AI reasoning capabilities.

A. Lawsen
arXiv:2506.09250 · cs.AI, cs.LG · submitted Jun 10, 2025 · updated Jun 16, 2025
abstract · pdf · html · Comment on: arXiv:2506.06941 Latest version removes Claude as a co-author, in line with arXiv policies, it also corrects mistakes in sections 4 and 6 of the original submission, as well as several typographical errors

add comment on HN
Also discussed: Jun 2025 (16 points, 14 comments) · Jun 2025 (4 points, 1 comment)

Doesn't seem to be the title, and yet the last submission has a similar title - so maybe it was renamed?

Title: Comment on The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity (which would need an edit to fit in 80 chars)

Discussion (15 points, 18 hours ago, 11 comments) https://news.ycombinator.com/item?id=44287172

Related:

An article with exactly this name (57 points, 8 days ago, 51 comments) https://news.ycombinator.com/item?id=44221900

"The Illusion of Thinking" – Thoughts on This Important Paper (54 points, 7 days ago, 75 comments) https://news.ycombinator.com/item?id=44234626

Original The Illusion of Thinking: Strengths and limitations of reasoning models (486 points, 10 days ago, 270 comments) https://news.ycombinator.com/item?id=44203562