about
True Detective: A Deep Abductive Reasoning Benchmark (arxiv.org)
8 points by optimalsolver on Apr 24, 2024 | hide | past | pdf | discuss on HN

In plain words: A new test asks AI to read 191 long detective stories and pick the correct solution from clues hidden in the text. The strongest AI solved only 38% of the puzzles, well below the average human solver.

Abstract · True Detective: A Deep Abductive Reasoning Benchmark Undoable for GPT-3 and Challenging for GPT-4

Large language models (LLMs) have demonstrated solid zero-shot reasoning capabilities, which is reflected in their performance on the current test tasks. This calls for a more challenging benchmark requiring highly advanced reasoning ability to be solved. In this paper, we introduce such a benchmark, consisting of 191 long-form (1200 words on average) mystery narratives constructed as detective puzzles. Puzzles are sourced from the "5 Minute Mystery" platform and include a multiple-choice question for evaluation. Only 47% of humans solve a puzzle successfully on average, while the best human solvers achieve over 80% success rate. We show that GPT-3 models barely outperform random on this benchmark (with 28% accuracy) while state-of-the-art GPT-4 solves only 38% of puzzles. This indicates that there is still a significant gap in the deep reasoning abilities of LLMs and humans and highlights the need for further research in this area. Our work introduces a challenging benchmark for future studies on reasoning in language models and contributes to a better understanding of the limits of LLMs' abilities.

Maksym Del, Mark Fishel
arXiv:2212.10114 · cs.CL · submitted Dec 20, 2022 · updated Jun 1, 2023
abstract · pdf · html · 5 pages, to appear at *SEM

add comment on HN