In plain words: A new test collects advanced reasoning problems in math, physics, biology, chemistry, and law, and grades each reasoning step against a checklist instead of just the final answer. Top models score far under 50% on the hardest tasks, and the checklist grading matched human judges.
Abstract · ARB: Advanced Reasoning Benchmark for Large Language Models
Large Language Models (LLMs) have demonstrated remarkable performance on various quantitative reasoning and knowledge benchmarks. However, many of these benchmarks are losing utility as LLMs get increasingly high scores, despite not yet reaching expert performance in these domains. We introduce ARB, a novel benchmark composed of advanced reasoning problems in multiple fields. ARB presents a more challenging test than prior benchmarks, featuring problems in mathematics, physics, biology, chemistry, and law. As a subset of ARB, we introduce a challenging set of math and physics problems which require advanced symbolic reasoning and domain knowledge. We evaluate recent models such as GPT-4 and Claude on ARB and demonstrate that current models score well below 50% on more demanding tasks. In order to improve both automatic and assisted evaluation capabilities, we introduce a rubric-based evaluation approach, allowing GPT-4 to score its own intermediate reasoning steps. Further, we conduct a human evaluation of the symbolic subset of ARB, finding promising agreement between annotators and GPT-4 rubric evaluation scores.
Tomohiro Sawada, Daniel Paleka, Alexander Havrilla, Pranav Tadepalli, Paula Vidas, Alexander Kranias, John J. Nay, Kshitij Gupta, Aran Komatsuzaki
arXiv:2307.13692 · cs.CL, cs.LG · submitted Jul 25, 2023 · updated Jul 28, 2023
abstract · pdf · html · Submitted to NeurIPS Datasets and Benchmarks Track
Researchers in child cognitive psychology make better tests for reasoning ability. Children have very limited knowledge, so the tests are based on core concepts like numbers, physical object properties like that object can't be in two places simultaneously. Also some objects have agency and some don't, some things are platonic and some concrete.
For example, can LLM can determine reliably what is a "moral subject", or "has agency" using simple language that is suitable for children (no specific terminology that works as a hint). Surely LLM can't do legal reasoning if it can't do that.
In my own experiments, LLM's are too much syntax driven to be reliable. When you frame the exact same question in 10 different ways, you get 4 different answers. As long as I can make LLM to think that Number 4 pencil is responsible for stabbing someone, it really can't do common sense reasoning.