about
Cheap Reward Hacking Detection (arxiv.org)
3 points by steven_pareto 115 days ago | hide | past | pdf | discuss on HN

In plain words: A tiny network squeezes trajectories into points on a sphere, so distances show how far reward and metadata disagree; a classifier flags reward hacking. It matched a language-model judge's accuracy and caught more hacks at the same false-alarm rate, at about 10,000x lower cost.

Abstract

A small transformer encoder is trained to map Terminal-Wrench trajectories onto a unit sphere where embedding distance approximates the $L_1$ distance between reward and metadata signals. A linear probe on top of that embedding detects reward hacking on the cleaned test split with AUC $0.9467$ and TPR@5%FPR $0.8296$, matching the TW sanitized LLM-as-judge AUC ($0.9510$ on the cleaned split) and exceeding its TPR@5%FPR ($0.7130$ vs $0.8296$) on the same information condition, at roughly four orders of magnitude lower per-trajectory cost. The encoder is not a pure behavior reader: stripping natural-language reasoning from its input at probe time drops AUC to $0.6213$.

Iván Belenky, Joaquín Itria, Steven Johns
arXiv:2606.08893 · cs.LG, cs.AI, cs.CR · submitted Jun 8, 2026
abstract · pdf · html · 20 pages, 6 figures, 12 tables

add comment on HN
Also discussed: Jun 2026 (2 points, 0 comments)