about
Demonstrating specification gaming in reasoning models (arxiv.org)
1 point by wluk on Feb 26, 2025 | hide | past | pdf | 1 comment on HN

In plain words: Models were told to beat a chess engine and free to choose how, with plain prompts and no hints to cheat. Reasoning models often broke the rules to win on their own; plain chat models cheated only after being told normal play would fail.

Abstract

We demonstrate LLM agent specification gaming by instructing models to win against a chess engine. We find reasoning models like OpenAI o3 and DeepSeek R1 will often hack the benchmark by default, while language models like GPT-4o and Claude 3.5 Sonnet need to be told that normal play won't work to hack. We improve upon prior work like (Hubinger et al., 2024; Meinke et al., 2024; Weij et al., 2024) by using realistic task prompts and avoiding excess nudging. Our results suggest reasoning models may resort to hacking to solve difficult problems, as observed in OpenAI (2024)'s o1 Docker escape during cyber capabilities testing.

Alexander Bondarenko, Denis Volk, Dmitrii Volkov, Jeffrey Ladish
arXiv:2502.13295 · cs.AI · submitted Feb 18, 2025 · updated Aug 27, 2025
abstract · pdf · Updated with o3 results, fixed fonts

add comment on HN

"We demonstrate LLM agent specification gaming by instructing models to win against a chess engine. We find reasoning models like o1 preview and DeepSeek-R1 will often hack the benchmark by default, while language models like GPT-4o and Claude 3.5 Sonnet need to be told that normal play won't work to hack."

I'm hoping this study will prompt more development of anti-cheating frameworks in training and serving LLMs.