about
Counting as a minimal probe of language model reliability (arxiv.org)
4 points by nateb2022 150 days ago | hide | past | pdf | discuss on HN

In plain words: Asking 126 AI models to count strings of repeated characters tests whether they can keep an exact state while applying one rule over and over. All failed abruptly past a limit; bigger models, more thinking time, and tools did not help.

Abstract · Language models fail at extended rule following

Large language models are highly capable of answering difficult questions by retrieving, recombining, and attending to information in long contexts. For agentic tasks, an additional capability is required: the preservation of an exact state while repeatedly applying rules. We find that this reliability is absent across language models. To demonstrate, we query 126 leading model variants with the task of counting a long string of repeated characters, and we find they all cannot accurately count above a model-dependent, syntax-sensitive counting capacity threshold. Failures are abrupt and persist even with increasing model size, inference time computation, and external tool. Mechanistic probing indicates that models use a finite number of internal states to mimic counting as a rule and fail once these states are exhausted. Furthermore, such states are the basis for performing complex tasks beyond counting. These results indicate that fundamentally new model architectures are required for autonomous agents to achieve truly reliable rule following capabilities.

Tianxiang Dai, Jonathan Fan
arXiv:2605.02028 · cs.CL · submitted May 3, 2026 · updated May 16, 2026
abstract · pdf · html · for accessing the data and code for reproduction of the study, see https://txdai.github.io/counting-reliability/

add comment on HN