about
Rethinking the Illusion of Thinking (arxiv.org)
2 points by jamesblonde on Jul 17, 2025 | hide | past | pdf | 1 comment on HN

In plain words: They retested two puzzle challenges—stacking disks and ferrying people across a river—with step-by-step hints and back-and-forth help. The disk puzzle still failed near 8 disks, but the river puzzle's failures were unsolvable setups; on solvable ones the models handled over 100 pairs.

Abstract

Earlier this year, Apple ignited controversy by publishing "The Illusion of Thinking," prompting heated debate within the AI community. Critics seized upon the findings as conclusive evidence that Large Reasoning Models (LRMs) lack genuine reasoning capabilities, branding them as mere stochastic parrots. Meanwhile, defenders-spearheaded by Lawsen et al. (2025)-fired back, condemning the experimental setup as flawed and the conclusions overstated. We clarify this debate by replicating and refining two of the original study's most contentious benchmarks: Towers of Hanoi and River Crossing. By introducing incremental stepwise prompting and agentic collaborative dialogue, we show that previously reported failures solving the Towers of Hanoi were not purely result of output constraints, but also partly a result of cognition limitations: LRMs still stumble when complexity rises moderately (around 8 disks). Moreover, the River Crossing results initially heralded as catastrophic failures turn out to hinge upon testing unsolvable configurations. Once we limit tests strictly to solvable problems-LRMs effortlessly solve large instances involving over 100 agent pairs. Our findings ultimately defy simplistic narratives: today's LRMs are stochastic, RL-tuned searchers in a discrete state space we barely understand. Real progress in symbolic, long-horizon reasoning demands mapping that terrain through fine-grained ablations like those introduced here.

Iñaki Dellibarda Varela, Pablo Romero-Sorozabal, Eduardo Rocon, Manuel Cebrian
arXiv:2507.01231 · cs.AI · submitted Jul 1, 2025
abstract · pdf · html · 8 pages, 4 figures

add comment on HN

In June 2025, Apple published a highly controversial paper, "The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity" that claimed Large Reasoning Models (LRM) did very little reasoning (planning).

Anthropic's Lawson fired back, condemning the experimental setup as flawed and the conclusions overstated.

This paper provides evidence supporting Apple's take - "failures solving the Towers of Hanoi were not purely result of output constraints, but also partly a result of cognition limitations: LRMs still stumble when complexity rises moderately (around 8 disks)"

" we also identified persistent failure modes that reveal limitations in long-horizon consistency and symbolic generalization. Our analysis suggests that these reasoning breakdowns stem not only from architectural constraints, but also from the inherently stochastic nature of these systems and the optimization methods they rely on."