In plain words: A new test checks whether AI models can handle data structures—like lists, trees, and networks—by asking them to perform operations on 4,140 problems covering 20 structures. The best model solved only 46% of the hardest ones, showing they struggle with structural reasoning.
Abstract
Large language models (LLMs) are deployed on increasingly complex tasks that require multi-step decision-making. Understanding their algorithmic reasoning abilities is therefore crucial. However, we lack a diagnostic benchmark for evaluating these capabilities. We propose to use data structures as a principled lens: as fundamental building blocks of algorithms, they naturally probe structural reasoning - the ability to understand and manipulate relationships such as order, hierarchy, and connectivity that underpin algorithmic reasoning. We introduce DSR-Bench (Data Structure Reasoning Benchmark), spanning 20 data structures, 35 operations, and 4,140 problem instances. DSR-Bench features hierarchical task organization, fully automated generation and evaluation, and fine-grained diagnostics. Evaluating 13 state-of-the-art LLMs reveals critical limitations: the top-performing model achieves only 0.46/1 on challenging instances. Three auxiliary probes targeting more realistic usages expose further weaknesses: models perform poorly on spatial data and context-rich scenarios, and they struggle to reason over their own code.
Yu He, Yingxi Li, Colin White, Ellen Vitercik
arXiv:2505.24069 · cs.LG, cs.AI · submitted May 29, 2025 · updated May 30, 2026
abstract · pdf · html · Proceedings of the 43rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026