about
ZebraLogic: On The Scaling Limits Of LLMs For Logical Reasoning (arxiv.org)
1 point by optimalsolver on Feb 10, 2025 | hide | past | pdf | 1 comment on HN

In plain words: A collection of classic logic grid puzzles with adjustable difficulty tests how well AI systems solve constraint puzzles as they get harder. Accuracy drops sharply as puzzles grow more complex, and bigger models or extra thinking time do not fix it.

Abstract · ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

We investigate the logical reasoning capabilities of large language models (LLMs) and their scalability in complex non-monotonic reasoning. To this end, we introduce ZebraLogic, a comprehensive evaluation framework for assessing LLM reasoning performance on logic grid puzzles derived from constraint satisfaction problems (CSPs). ZebraLogic enables the generation of puzzles with controllable and quantifiable complexity, facilitating a systematic study of the scaling limits of models such as Llama, o1 models, and DeepSeek-R1. By encompassing a broad range of search space complexities and diverse logical constraints, ZebraLogic provides a structured environment to evaluate reasoning under increasing difficulty. Our results reveal a significant decline in accuracy as problem complexity grows -- a phenomenon we term the curse of complexity. This limitation persists even with larger models and increased inference-time computation, suggesting inherent constraints in current LLM reasoning capabilities. Additionally, we explore strategies to enhance logical reasoning, including Best-of-N sampling, backtracking mechanisms, and self-verification prompts. Our findings offer critical insights into the scalability of LLM reasoning, highlight fundamental limitations, and outline potential directions for improvement.

Bill Yuchen Lin, Ronan Le Bras, Kyle Richardson, Ashish Sabharwal, Radha Poovendran, Peter Clark, Yejin Choi
arXiv:2502.01100 · cs.AI, cs.CL, cs.LG · submitted Feb 3, 2025 · updated Jul 15, 2025
abstract · pdf · html · Accepted to ICML 2025

add comment on HN

Thanks for posting this. I had just ran across this one and wondered if I should post it myself. Came here to see if someone had. I'd gone looking for this after reading about DeepScaleR yesterday.

I'm excited about the prospect of using formal methods to generate, analytically (not via LLM), synthetic problem+solution sets of perfect quality, of progressive size & complexity, for training; and then using RL, with progressive scaling, as in DeepScaleR.