about
Evaluating Reasoning by LLMs Using the New York Times Connections Word Game (arxiv.org)
2 points by PaulHoule on Jun 26, 2024 | hide | past | pdf | discuss on HN

In plain words: Large language models were tested on 438 New York Times Connections puzzles, where players sort 16 words into four hidden groups, and scored against novice and expert humans. The best model fully solved just 18% of games, losing to both human groups.

Abstract · Connecting the Dots: Evaluating Abstract Reasoning Capabilities of LLMs Using the New York Times Connections Word Game

The New York Times Connections game has emerged as a popular and challenging pursuit for word puzzle enthusiasts. We collect 438 Connections games to evaluate the performance of state-of-the-art large language models (LLMs) against expert and novice human players. Our results show that even the best performing LLM, Claude 3.5 Sonnet, which has otherwise shown impressive reasoning abilities on a wide variety of benchmarks, can only fully solve 18% of the games. Novice and expert players perform better than Claude 3.5 Sonnet, with expert human players significantly outperforming it. We create a taxonomy of the knowledge types required to successfully cluster and categorize words in the Connections game. We find that while LLMs perform relatively well on categorizing words based on semantic relations they struggle with other types of knowledge such as Encyclopedic Knowledge, Multiword Expressions or knowledge that combines both Word Form and Meaning. Our results establish the New York Times Connections game as a challenging benchmark for evaluating abstract reasoning capabilities in AI systems.

Prisha Samadarshi, Mariam Mustafa, Anushka Kulkarni, Raven Rothkopf, Tuhin Chakrabarty, Smaranda Muresan
arXiv:2406.11012 · cs.CL, cs.AI · submitted Jun 16, 2024 · updated Oct 14, 2024
abstract · pdf · html

add comment on HN