about
I Am a Strange Dataset: Metalinguistic Tests for Language Models (arxiv.org)
1 point by birriel on Jan 11, 2024 | hide | past | pdf | discuss on HN

In plain words: A hand-built set of tricky sentences that talk about themselves, like "The penultimate word in this sentence is," asks models to finish them or judge whether they are true. Every model tested scored near chance, while untrained people hit 89–93%.

Abstract · I am a Strange Dataset: Metalinguistic Tests for Language Models

Statements involving metalinguistic self-reference ("This paper has six sections.") are prevalent in many domains. Can current large language models (LLMs) handle such language? In this paper, we present "I am a Strange Dataset", a new dataset for addressing this question. There are two subtasks: generation and verification. In generation, models continue statements like "The penultimate word in this sentence is" (where a correct continuation is "is"). In verification, models judge the truth of statements like "The penultimate word in this sentence is sentence." (false). We also provide minimally different metalinguistic non-self-reference examples to complement the main dataset by probing for whether models can handle metalinguistic language at all. The dataset is hand-crafted by experts and validated by non-expert annotators. We test a variety of open-source LLMs (7B to 70B parameters) as well as closed-source LLMs through APIs. All models perform close to chance across both subtasks and even on the non-self-referential metalinguistic control data, though we find some steady improvement with model scale. GPT 4 is the only model to consistently do significantly better than chance, and it is still only in the 60% range, while our untrained human annotators score well in the 89-93% range. The dataset and evaluation toolkit are available at https://github.com/TristanThrush/i-am-a-strange-dataset.

Tristan Thrush, Jared Moore, Miguel Monares, Christopher Potts, Douwe Kiela
arXiv:2401.05300 · cs.CL, cs.AI · submitted Jan 10, 2024 · updated Aug 6, 2024
abstract · pdf · html · ACL 2024

add comment on HN