In plain words: Questions are paired up and given more answer choices to make old tests harder again. Stronger models lose accuracy in a predictable way, raising the score ceiling so benchmarks everyone now aces become useful again.
Abstract
Recent work showed that small changes in benchmark questions can reduce LLMs' reasoning and recall. We explore two such changes: pairing questions and adding more answer options, on three benchmarks: WMDP-bio, GPQA, and MMLU variants. We find that for more capable models, these predictably reduce performance, essentially heightening the performance ceiling of a benchmark and unsaturating it again. We suggest this approach can resurrect old benchmarks.
Igor Ivanov, Dmitrii Volkov
arXiv:2502.06738 · cs.LG · submitted Feb 10, 2025
abstract · pdf · html